Leaderboard
Legal Benchmarks on real-world legal work.
- Substance: Whether the legal work is correct, complete, and responsive. Measured as task pass rate, the percentage of tasks for which every applicable criterion was satisfied.
- Form (1 to 3): Whether a lawyer would want to use the delivered work as written, including clarity, structure, and right-sizing.
Results, contract workflows
| Claude Opus 4.8 | Anthropic | 67.6% | 2.67 | ~$0.29 |
| Claude Fable 5 | Anthropic | 61.8% | 2.66 | ~$0.64 |
| Claude Opus 5New | Anthropic | 61.8% | 2.47 | ~$0.74 |
| Grok 4.5New | xAI | 58.8% | 2.61 | ~$0.19 |
| Kimi K3New | Moonshot | 58.8% | 2.64 | ~$0.14 |
| Gemini 3.5 Flash | 55.9% | 2.60 | ~$0.08 | |
| Claude Sonnet 4.6 | Anthropic | 50.0% | 2.63 | $0.13 |
| Gemini 3.1 Pro | 50.0% | 2.69 | $0.07 | |
| GPT 5.6 SolNew | OpenAI | 44.1% | 2.75 | ~$0.19 |
| Qwen3.7 Max | Alibaba | 44.1% | 2.67 | ~$0.03 |
| GPT-5.5 | OpenAI | 41.2% | 2.77 | $0.15 |
| DeepSeek V4 Pro | DeepSeek | 26.5% | 2.68 | ~$0.03 |
| GPT-5.4-mini | OpenAI | 26.5% | 2.55 | ~$0.01 |
Key takeaways
- Anthropic still sets the pace in drafting, but its newest model does not lead. Claude Opus 4.8 remains first. Opus 5 finishes behind it and ties Fable 5, so a newer model name does not guarantee better drafts.
- Token pricing alone is misleading. The cost of completing the task is what matters. Opus 5 costs less per token than Fable 5, but it processes more input and writes much longer answers. A finished task costs about $0.74 versus $0.64.
- Kimi K3 is the strongest drafter outside Anthropic. At roughly $0.14 per task, it drafts nearly as well as the leaders and outperforms every Google and OpenAI model tested.
- GPT 5.6 Sol's drafts are more polished than they are accurate. The writing looks clean and final, but more than half of its drafts miss at least one instruction. Review the substance, not just the polish.
- Conflict detection remains an open weakness across the field. Most models keep drafting when instructions clash with the documents or with each other. Opus 4.8 handles these cases best, while GPT 5.6 Sol is the most likely to continue without raising the problem.
How we test
Every model receives the same instructions and native source documents of lawyer-authored tasks. A task passes only when every applicable criterion is satisfied; form is scored separately for clarity, length, and structure. LLM graders are checked through agreement analysis, human spot reviews, and provider-favoritism tests.
Published research
Real benchmarks, real results
- FrameworkAugust 2026
Methodology
How Legal Benchmarks evaluates AI on real-world legal work: what we test, how we run models and applications, how we score their output, and how we validate the graders behind both leaderboards.
- ReportUpdated quarterly
The Leaderboard
AI systems tested on real-world legal work.