Canada's open-weight model lab. We train, quantize, and deploy sovereign AI on Canadian Blackwell silicon — for the regulated industries that can't run on someone else's API.
Public capital is flowing through every layer of the Canadian AI stack — data centres, compute subsidies, defence platforms, sovereign cloud. The one layer that's empty is the one where the IP actually lives.
Canada is investing $2 billion in sovereign AI. No open, sovereign model lab exists.
Every dollar today flows through closed APIs, telco-hosted metal, or thin wrapper layers. Cohere just open-weighted Command A+ — a Canadian foundation worth building on. The layer above it — vertical-specialized, audited, customer-owned — is still unoccupied.
We're not a wrapper. Not a hosting platform. Not a fine-tuning API. We take frontier open base models and produce shippable, quantized, sovereign deployments — with the IP and audit evidence customers actually own.
Post-training on open base models. SFT, DPO, GRPO, RLAIF. Customer data stays sovereign throughout the run.
Production W4A16, NVFP4, MXFP4 recipes — with MTP draft heads preserved so speculative decoding survives end-to-end. Patches land upstream in vLLM and llm-compressor.
Air-gapped, on-prem, sovereign cloud. With audit trails, eval evidence, and model risk documentation packaged with the build.
We're starting wide. The verticals will narrow themselves as contracts land — not as we pretend to know. Each one has a clear buyer set, a clear corpus, and a clear wedge against frontier general models.
Hallucination floors below frontier models on Canadian case law.
On-prem deployable, PHIPA/PIPEDA-clean, evidence-cited outputs.
Air-gapped, classified-ready, doctrine-aware, Five Eyes interoperable.
OSFI-aligned, MRM-documented, audit-ready, on Canadian soil.
Sovereign AI in Canada has shifted from policy paper to active deployment in under twelve months. The first lab to ship audited, vertical, open-weight models with defence contracts behind them owns the category.
Engineering notes and release writeups from the lab — recipes, evaluation methodology, and the walls we hit along the way. The same rigor that goes into the model cards, in long form.
We quantized Tencent's Hy3 (295B MoE) to 4-bit weights while keeping its MTP draft layer in BF16. It matches the FP8 release on quality at 57% of the footprint, runs +40% faster single-stream than the same-scheme quant that shipped without MTP, and takes the high-concurrency throughput crown among 4-bit Hy3 quants.
We quantized GLM-5.2 (744B MoE) to 4-bit weights while keeping its MTP draft head in BF16. It matches the FP8 release on quality, fits on four H200s instead of eight, and is the fastest popular 4-bit GLM-5.2 quant in the interactive serving regime.
If your organization needs a model it can audit, deploy on-prem, and keep under Canadian jurisdiction — we should talk. We respond to every serious inquiry within 48 hours.
Or email partnerships@cql.ca · press@cql.ca directly.