CanLegal Bench is live: a public leaderboard for how well AI reasons about Canadian law

Our Canadian legal-reasoning benchmark has gone from a design doc to a live public leaderboard: 14,547 scored records from primary legal sources, 17 task types, 14 jurisdictions, both official languages and both legal traditions. Five frontier models, closed-book. The top three sit within three points of each other, and none of them clears 60% on the four hardest tasks.

TL;DR — CanLegal Bench is now public at canlegalbench.com, with a live leaderboard. It scores models on 14,547 primary-source records across 17 task types and 14 jurisdictions; 31% of records are in French, and the suite spans common-law and civil-law material. Each model answers closed-book: no tools, no retrieval, no web. In the v0.8.3 snapshot on 8 Oct 2026, Claude Fable 5 leads at 80.1, ahead of GPT-5.6 Sol (78.1) and Gemini 3.1 Pro (77.5), followed by Grok 4.5 (71.7) and Cohere Command A+ (49.5). No model clears 60% on citation resolution, citation holdings, bijural equivalence or issue spotting. Next: the full interactive leaderboard, then the open evaluation release (data, harness and scoring code), so anyone can reproduce every score.

Why we built it

Legal is one of our four verticals, and our pitch there is specific: lower hallucination rates than frontier models on Canadian case law. We can only make that claim if we can measure it, and nothing measured it. General legal benchmarks are mostly American. They don't test Quebec civil law next to common-law provinces, they don't check that an answer holds in both English and French, and they don't ask whether a model can tell a real neutral citation from a fabricated one. That last skill matters most to a lawyer filing a factum.

When we listed CanLegal Bench on our models page it was marked Designing. Today it is a running benchmark with a public leaderboard, its own site, and a release plan.

What it measures

17 task types, chosen to match the work a Canadian legal professional actually does:

15 of the 17 task types include French-language records. Reference answers come from statutes, regulations, and court and tribunal decisions. They are checked by automated validation gates, two benchmark audits and an AI council that verified the answer keys against the primary sources. Records found defective are quarantined and excluded from scoring.

How it is scored

The leaderboard today

CanLegal Bench v0.8.3, snapshot taken 8 Oct 2026 from the live leaderboard. Suite mean = unweighted mean of strict accuracy over 17 task types, 14,547 records per model, closed-book. 95% CI in brackets. The live board at canlegalbench.com is authoritative.
ModelSuite mean95% CI
Claude Fable 580.178.9–81.4
GPT-5.6 Sol78.177.0–79.3
Gemini 3.1 Pro77.576.4–78.6
Grok 4.571.770.4–72.9
Cohere Command A+49.548.0–51.1

Claude Fable 5 leads by 2.0 points, but its interval slightly overlaps GPT-5.6 Sol's, so first place is not yet settled beyond the noise. Gemini 3.1 Pro is statistically level with GPT-5.6 Sol. Grok 4.5 is clearly fourth. Cohere Command A+, the only Canadian-built model on the board, finishes 22 points behind the next model.

Where models break

The suite mean hides a split. On some tasks the frontier is close to solved:

Four tasks defeat every model:

Strict accuracy (%), CanLegal Bench v0.8.3, 8 Oct 2026. The best score on each of these tasks is below 60%.
TaskRecordsBestRange across five models
Citation Resolution1,50044.40.0–44.4
Citation Holding1,39556.11.1–56.1
Bijural Equivalence17557.133.1–57.1
Issue Spotting43858.739.3–58.7

The pattern is the one practitioners worry about. Models are good at confirming that a citation exists and poor at saying what the case actually held. Statute QA shows the widest spread on the board: 81.2% for the leader, 64.8% for second place, 9.9% at the bottom. Regulation QA tops out at 62.9%. Detailed recall of Canadian legislation still separates models, and nobody has it nailed.

What changed recently

Limitations (honestly)

What's next

To hear when those ship, use Notify me on the benchmark site. Questions and partnerships: partnerships@cql.ca.

See the leaderboard

canlegalbench.com — live results, task descriptions and release updates for CanLegal Bench, developed by Canada Quant Labs. Snapshot in this post: v0.8.3, 8 Oct 2026.