Skip to content

Leaderboard

Legal Benchmarks on real-world legal work.
  • Substance: Whether the legal work is correct, complete, and responsive. Measured as task pass rate, the percentage of tasks for which every applicable criterion was satisfied.
  • Form (1 to 3): Whether a lawyer would want to use the delivered work as written, including clarity, structure, and right-sizing.

Results, contract workflows

Contract Workflows results: task pass rate and form for each model, plus cost per task.
Claude Opus 4.8Anthropic67.6%2.67~$0.29
Claude Fable 5Anthropic61.8%2.66~$0.64
Claude Opus 5NewAnthropic61.8%2.47~$0.74
Grok 4.5NewxAI58.8%2.61~$0.19
Kimi K3NewMoonshot58.8%2.64~$0.14
Gemini 3.5 FlashGoogleGoogle55.9%2.60~$0.08
Claude Sonnet 4.6Anthropic50.0%2.63$0.13
Gemini 3.1 ProGoogleGoogle50.0%2.69$0.07
GPT 5.6 SolNewOpenAI44.1%2.75~$0.19
Qwen3.7 MaxAlibaba44.1%2.67~$0.03
GPT-5.5OpenAI41.2%2.77$0.15
DeepSeek V4 ProDeepSeek26.5%2.68~$0.03
GPT-5.4-miniOpenAI26.5%2.55~$0.01

Key takeaways

  • Anthropic still sets the pace in drafting, but its newest model does not lead. Claude Opus 4.8 remains first. Opus 5 finishes behind it and ties Fable 5, so a newer model name does not guarantee better drafts.
  • Token pricing alone is misleading. The cost of completing the task is what matters. Opus 5 costs less per token than Fable 5, but it processes more input and writes much longer answers. A finished task costs about $0.74 versus $0.64.
  • Kimi K3 is the strongest drafter outside Anthropic. At roughly $0.14 per task, it drafts nearly as well as the leaders and outperforms every Google and OpenAI model tested.
  • GPT 5.6 Sol's drafts are more polished than they are accurate. The writing looks clean and final, but more than half of its drafts miss at least one instruction. Review the substance, not just the polish.
  • Conflict detection remains an open weakness across the field. Most models keep drafting when instructions clash with the documents or with each other. Opus 4.8 handles these cases best, while GPT 5.6 Sol is the most likely to continue without raising the problem.

How we test

Every model receives the same instructions and native source documents of lawyer-authored tasks. A task passes only when every applicable criterion is satisfied; form is scored separately for clarity, length, and structure. LLM graders are checked through agreement analysis, human spot reviews, and provider-favoritism tests.

Get benchmarked for the next quarterly awards