The short version — Canada Quant Labs is hiring an AI Engineer to own pieces of our train → quantize → deploy pipeline. You'll ship open-weight checkpoints (like our Hy3, GLM-5.2, and DeepSeek-V4 quants), build the eval harnesses that prove they hold up, and serve them through vLLM. Full-time, hybrid in Victoria, BC. CAD $140,000–$190,000 + equity.
The role
We're Canada's open-weight model lab: we train, quantize, and deploy sovereign AI on Canadian Blackwell silicon for the industries — legal, medical, defence, finance — that can't run on someone else's API. Everything we ship is public: weights, recipes, benchmarks, and the engineering log of what broke along the way.
This is a hands-on engineering role on the model team, not a research-caretaker role. The person in it trains models, breaks them, measures them, and ships them — then writes up what happened so the community doesn't have to re-walk the dead-ends.
What you'll do
- Train and fine-tune open-weight models on our DGX B300 cluster — domain distills, reasoning and agentic post-training, and Canadian-corpora adaptations for legal, medical, defence, and finance buyers.
- Own quantization recipes end-to-end — W4A16 and NVFP4 with
llm-compressor, calibration-data selection, and quality validation against the FP8/BF16 references. - Build and run the eval harnesses that decide what ships: reasoning, instruction-following, long-context retrieval, and domain benchmarks like our in-progress CanLegal-Bench.
- Serve what you ship through vLLM — speculative decoding (MTP), expert-parallel MoE serving, throughput/latency profiling, and the kernel-level debugging when numbers don't make sense.
- Publish everything — model cards, reproducible recipes, benchmark methodology, and engineering logs — and upstream fixes to the open stack (vLLM, llm-compressor) when you find them.
What you bring
- Strong PyTorch and transformer fundamentals. You've trained or fine-tuned multi-billion-parameter models and can reason about what's happening inside the network — not just drive the framework.
- Real LLM serving or quantization experience. You've shipped something with vLLM, TensorRT-LLM, SGLang, or similar — or you've built quant/serving tooling (GPTQ, AWQ, NVFP4, activation-aware recipes) that other people used.
- Systems instincts. Comfortable reading CUDA/perf profiles, reasoning about memory footprints and tensor/expert parallelism, and hunting a 13% throughput regression down to a kernel flag.
- Evidence of craft. A checkpoint, a kernel, an eval harness, a merged PR, a write-up — something public we can look at. We hire from artifacts, not credentials.
- Clear technical writing. Our work is open by default; your model cards and engineering logs are part of the product.
Nice to have
- MoE training or serving experience (expert parallelism, router behaviour, sparse-attention variants).
- Speculative decoding — MTP/EAGLE-style draft heads, acceptance-rate tuning.
- Eval science: benchmark design, contamination hygiene, statistical rigour about run-to-run noise.
- Background in a regulated domain (legal, medical, defence, finance) or with Canadian data-residency constraints.
- Open-source maintainership or a meaningful contribution history in the inference/training stack.
How we work
Small team, hybrid in Victoria, BC — we like being in the room for design reviews and launch days, and we don't count hours the rest of the time. Compute is dedicated, not queued: experiments start when you're ready. Decisions get written down, results get published, and the person who did the work presents it.
Compensation & logistics
| Base salary | CAD $140,000–$190,000 |
| Equity | Meaningful early-stage grant |
| Location | Victoria, BC · hybrid (relocation support available) |
| Work authorization | Canadian citizens & permanent residents preferred |
| Start | Flexible · we're hiring now |
How to apply
Apply below — a short note and your résumé are all we need. Links to artifacts (Hugging Face, GitHub, papers, benchmarks, write-ups) help a lot.
We read every application and reply within two weeks. The process: intro call → technical deep-dive on something you've built → paid half-day work trial on real cluster work → offer.