TL;DR — CanLegal Bench is now public at canlegalbench.com, with a live leaderboard. It scores models on 14,547 primary-source records across 17 task types and 14 jurisdictions; 31% of records are in French, and the suite spans common-law and civil-law material. Each model answers closed-book: no tools, no retrieval, no web. In the v0.8.3 snapshot on 8 Oct 2026, Claude Fable 5 leads at 80.1, ahead of GPT-5.6 Sol (78.1) and Gemini 3.1 Pro (77.5), followed by Grok 4.5 (71.7) and Cohere Command A+ (49.5). No model clears 60% on citation resolution, citation holdings, bijural equivalence or issue spotting. Next: the full interactive leaderboard, then the open evaluation release (data, harness and scoring code), so anyone can reproduce every score.
Why we built it
Legal is one of our four verticals, and our pitch there is specific: lower hallucination rates than frontier models on Canadian case law. We can only make that claim if we can measure it, and nothing measured it. General legal benchmarks are mostly American. They don't test Quebec civil law next to common-law provinces, they don't check that an answer holds in both English and French, and they don't ask whether a model can tell a real neutral citation from a fabricated one. That last skill matters most to a lawyer filing a factum.
When we listed CanLegal Bench on our models page it was marked Designing. Today it is a running benchmark with a public leaderboard, its own site, and a release plan.
What it measures
17 task types, chosen to match the work a Canadian legal professional actually does:
- Citations: Cite Verify, Citation Lookup, Citation Holding, Citation Resolution — does the case exist, what does the citation point to, and what did it decide?
- Statutes and regulations: Statute QA, Regulation QA, Limitation Periods, Temporal Reasoning — reading legislation correctly and applying it to dates and deadlines.
- Case law: Case Outcome, Case Treatment, Standard of Review, Doctrine Shift, Tribunal QA — predicting and characterizing how courts and tribunals rule.
- Legal analysis: Issue Spotting, Multi-Doc Reasoning, Jurisdiction Router — finding the issues in a fact pattern, reasoning across several sources, and sending a question to the right forum.
- Bijural Equivalence: matching legal concepts across Canada's common-law and civil-law traditions.
15 of the 17 task types include French-language records. Reference answers come from statutes, regulations, and court and tribunal decisions. They are checked by automated validation gates, two benchmark audits and an AI council that verified the answer keys against the primary sources. Records found defective are quarantined and excluded from scoring.
How it is scored
- Closed-book. No tools, retrieval or web access during evaluation. The scores measure what a model knows, not what it can look up.
- Same prompts, same frozen record set for every model.
- Strict accuracy on judge-scored tasks. An answer counts only when the judge gives it full credit. For issue spotting, every key issue must be identified; extra issues are not penalized.
- Declining is not a free pass. On Citation Resolution, a declined answer counts as not correct.
- Suite mean is unweighted across the 17 task types, so the 7,911-record Citation Lookup task counts no more than the 57-record Jurisdiction Router. Each mean is reported with a 95% confidence interval.
- Perturbation and contamination safeguards test whether a model is reasoning or pattern-matching against memorized text.
The leaderboard today
| Model | Suite mean | 95% CI |
|---|---|---|
| Claude Fable 5 | 80.1 | 78.9–81.4 |
| GPT-5.6 Sol | 78.1 | 77.0–79.3 |
| Gemini 3.1 Pro | 77.5 | 76.4–78.6 |
| Grok 4.5 | 71.7 | 70.4–72.9 |
| Cohere Command A+ | 49.5 | 48.0–51.1 |
Claude Fable 5 leads by 2.0 points, but its interval slightly overlaps GPT-5.6 Sol's, so first place is not yet settled beyond the noise. Gemini 3.1 Pro is statistically level with GPT-5.6 Sol. Grok 4.5 is clearly fourth. Cohere Command A+, the only Canadian-built model on the board, finishes 22 points behind the next model.
Where models break
The suite mean hides a split. On some tasks the frontier is close to solved:
- Limitation Periods: 95.3–100% for the top four.
- Standard of Review: 96.8–98.9% for the top four.
- Cite Verify (is this neutral citation real?): 93.3–97.8% for the top four.
Four tasks defeat every model:
| Task | Records | Best | Range across five models |
|---|---|---|---|
| Citation Resolution | 1,500 | 44.4 | 0.0–44.4 |
| Citation Holding | 1,395 | 56.1 | 1.1–56.1 |
| Bijural Equivalence | 175 | 57.1 | 33.1–57.1 |
| Issue Spotting | 438 | 58.7 | 39.3–58.7 |
The pattern is the one practitioners worry about. Models are good at confirming that a citation exists and poor at saying what the case actually held. Statute QA shows the widest spread on the board: 81.2% for the leader, 64.8% for second place, 9.9% at the bottom. Regulation QA tops out at 62.9%. Detailed recall of Canadian legislation still separates models, and nobody has it nailed.
What changed recently
- Citation Resolution launched on 4 Oct 2026 and joined the suite mean on 5 Oct, once every model on the board had a score for it. It is now the hardest task on the board.
- Three models retired on 30 Sep 2026: Claude Opus 4.8, GPT-5.5 and Grok 4.3. Newer versions had superseded them, and their cells were scored on older answer keys and stale generations. We retired them rather than mix old and new scoring on one board.
Limitations (honestly)
- Not yet reproducible by you. Until the open evaluation release ships, the data, harness and scoring code are not public, so these numbers can't be independently checked. Closing that gap is the next milestone.
- Closed-book only. The board measures what models know unaided. A retrieval-backed legal tool would score differently, and we don't measure that here.
- Judge-scored tasks depend on an LLM judge. Strict full-credit scoring keeps partial answers from being rounded up, but it is still a judge.
- Small tasks are noisy. Jurisdiction Router (57 records) and Doctrine Shift (70) are small enough that per-task gaps of a few points between models are within the noise. The suite-mean intervals don't capture that per-task uncertainty.
- Snapshot. The tables above are fixed as of 8 Oct 2026. The leaderboard is live, so check it for current numbers.
What's next
- Full interactive leaderboard on canlegalbench.com.
- Open evaluation release: the benchmark data, harness and scoring code, so anyone can reproduce every score.
To hear when those ship, use Notify me on the benchmark site. Questions and partnerships: partnerships@cql.ca.