Skip to content

Model Leaderboard

An independent comparison of how leading AI models perform on Contract Workflows and Data Extraction.
  • Substance: Whether the legal work is correct, complete, and responsive. Task pass is the share of tasks that satisfy every applicable criterion in both of the model's independent attempts.
  • Form: whether a lawyer would want to use the delivered work as written. Measured as the percentage of applicable binary form checks passed in both of the model's independent attempts.

Contract workflow leaderboard

Contract Workflows results: task pass rate and form for each model, plus cost per task and latency.
Rank
1Claude Opus 5.5Anthropic48.4%91.7%93.9%$1.473m 03s
1Claude Fable 5.1Anthropic48.4%90.7%89.7%$1.503m 51s
1Muse Spark 1.3Meta48.4%87.4%87.8%$0.204m 35s
4Claude Opus 5Anthropic45.2%90.1%90.8%$0.904m 02s
5Qwen 3.8 FlashQwen38.7%49.1%67.7%$0.038m 21s
6Claude Sonnet 5.5Anthropic29.0%85.3%91.1%$0.981m 54s
6Qwen 3.8 MaxQwen29.0%32.2%64.2%$0.396m 21s
8GPT-6 AstraOpenAI25.8%87.1%91.9%$1.132m 55s
8Gemini 3.8 FlashGoogle25.8%65.9%76.4%$0.226m 11s
8Grok 4.6xAI25.8%62.2%76.8%$0.284m 28s
11GPT-5.6 SolOpenAI22.6%84.8%89.2%$0.182m 51s
11Kimi K3Moonshot AI22.6%78.8%84.7%$0.262m 10s
11DeepSeek V4 FlashDeepSeek22.6%57.1%73.4%$0.014m 18s
11GLM-5.3 FlashZ.ai22.6%56.0%67.8%$0.024m 21s
15GPT-5.6 LunaOpenAI19.4%80.9%86.1%$0.022m 54s
15DeepSeek V4 ProDeepSeek19.4%78.6%88.3%$0.083m 58s
15GLM-5.3Z.ai19.4%68.1%81.9%$0.093m 31s
18GPT-5.6 TerraOpenAI16.1%81.2%86.0%$0.141m 41s
18Claude Sonnet 5Anthropic16.1%80.1%86.9%$0.362m 52s
20MiniMax M3MiniMax12.9%80.1%83.4%$0.032m 03s
21Mistral Medium 3.5Mistral9.7%47.7%75.0%$0.322m 52s

Key takeaways

  • The drafting lead is shared, at very different prices. Claude Opus 5.5, Claude Fable 5.1, and Muse Spark 1.3 pass the same share of contract tasks, while Muse Spark costs about 86% less than either Claude model. A top score no longer implies a top price.
  • Meeting almost every criterion is not the same as finishing the task. GPT-6 Astra meets 87.1% of drafting criteria yet completes about a quarter of tasks end to end. One missed requirement fails the whole task, so review against the full instruction, not overall polish.
  • A high task pass rate can rest on narrow coverage. Qwen 3.8 Flash sits near the top on completed drafting tasks while meeting fewer than half of the criteria overall. Its failures cluster in a few tasks, so check the criteria pass rate before trusting a rank.
  • Still no model to leave drafting unsupervised. The best contract result leaves more than half of workflows short of at least one requirement. Keep lawyer review in the loop.

Methodology

What we benchmark

The Models leaderboard evaluates AI models directly through the controlled Legal Benchmarks Harness.

It measures a model’s ability to interpret instructions, analyse legal materials, and complete real-world legal work under standardized conditions. It does not evaluate the surrounding application, interface, or proprietary document-handling system through which a customer might ordinarily access the model.

Every model is evaluated by Legal Benchmarks under our control. Results are never self-reported.

The task set

Models complete demanding, real-world legal assignments contributed by practitioners in the Legal Benchmarks community. The set covers two principal categories:

  • Contract workflows: producing or amending contract language in response to a lawyer’s instructions, including basic clause drafting, template-based drafting, and bespoke drafting.
  • Data extraction: locating and reporting information from individual documents or sets of related documents, including imperfect source materials such as scans, images, and files containing tracked changes.

Practitioner-authored tasks make up the core of the set. Synthetic tasks supplement them by targeting failure modes identified in earlier Legal Benchmarks research and published legal-AI benchmarking work. Every synthetic task is validated by a human expert before entering the benchmark.

Each task has a fixed set of binary pass/fail criteria authored and reviewed by lawyers. These criteria specify what a satisfactory answer must include and avoid while accommodating different approaches that would be professionally defensible.

How models are evaluated

Each task is completed in a sealed workspace in the Legal Benchmarks Harness.

The model receives the lawyer’s instruction and source documents, works independently using a small, fixed set of file tools, and submits a final answer together with any documents it produces.

Every model completes each task twice, in two independent attempts. The attempts are isolated from each other, and published scores require the work to succeed in both.

Evaluation conditions are standardized:

  • Every model receives the same task instructions and tools.
  • Documents are accessed through the same shared document-reading system.
  • The workspace has no internet access.
  • Every attempt starts fresh, with no memory of another task or attempt.
  • Each attempt has a fixed 15-minute limit and a fixed tool-call budget.
  • A model that does not finish within those limits fails the attempt.

Where a provider offers a reasoning-effort control, we request the high setting. Where no such control is available, the provider’s default settings apply.

What we measure

Model output quality is reported on two independent axes: substance and form. These scores are kept separate because a polished answer may still be legally unreliable, while a substantively correct answer may require editing before use.

Substance

Substance measures whether the legal work is correct, complete, responsive, and appropriately candid about uncertainty.

We report two substance measures:

  • Criteria pass rate: the percentage of applicable lawyer-authored criteria satisfied in both of the model’s attempts.
  • Task pass rate: the percentage of tasks that satisfy every applicable substance criterion in both attempts.

Task pass rate is the headline score for the Models leaderboard. It is intentionally strict: a task passes only if the model fails no applicable substance criterion in either attempt.

Form

Form measures whether the output is usable as delivered.

Each submission is assessed against a separate set of binary form checks. The form score is the percentage of applicable checks passed in both attempts.

Form does not affect the substance score or ranking.

Grading

Judges assess the model’s final submission, including its written answer and any files delivered. They do not grade the model’s internal reasoning or tool-use process.

Substance and form are graded independently. The substance judge does not see the form scores or comments, and the form judge does not see the substance assessment.

Substance is graded one criterion at a time by GLM 5.3 Flash. Each call determines whether the submission passes or fails one fixed lawyer-authored criterion.

Form is graded using the same model with a separate checklist, prompt, and judge configuration.

Judge calibration

The substance judge is calibrated against lawyer PASS/FAIL decisions for individual criteria in legal AI application outputs from the same task set.

Calibration cases focus on outputs where judges disagreed or where the answer was close to passing. Each case has one lawyer label.

Candidate judge configurations are compared using balanced accuracy: the mean of the judge’s PASS recall and FAIL recall. The prompt is selected using development tasks and then evaluated on separate tasks that were held out from prompt selection.

This calibration tests criterion-level agreement with lawyers on the selected sample. It does not establish the judge’s accuracy across every leaderboard output or its agreement with a lawyer’s overall assessment of an answer.

The model form judge does not currently have a lawyer-labelled accuracy measurement.

Results and rankings

Results are reported separately for Contract Workflows and Data Extraction, alongside overall figures.

Models are ranked by substance task pass rate. Tied models share a rank and are ordered by criteria pass rate. Criteria pass rate and form are reported separately to show:

  • Overall coverage of the legal requirements
  • Quality of presentation

Rankings represent a point-in-time evaluation and are refreshed whenever the benchmark is rerun. Cost is an average per attempt within each workflow; time is a model-level average across the full benchmark. Neither affects rank.

Model and application scores are not directly comparable. Models use shared document-reading tools and binary form checks, while applications are evaluated through their own interfaces and document-handling systems.

Benchmark security

The task instructions, reference documents, answer keys, and per-task criteria are not publicly distributed.

Releasing them would allow models and providers to train or tune directly against the benchmark, weakening its value as an independent evaluation. Researchers and academics may contact Legal Benchmarks to request further access to the methodology.

Limitations

  • Coverage: The task set is English only and currently leans towards commercial, IP, employment, M&A, and competition work, with a US and UK concentration.
  • Single assignments: Each task consists of one instruction and one submission. The benchmark does not yet test iterative refinement, multi-turn workflows, or longer matters.
  • Point-in-time results: Models and providers change frequently. A result from one evaluation period may not represent a later model version.
  • LLM judging: Judge calibration measures criterion-level agreement on a selected lawyer-labelled sample. It does not establish accuracy across the full leaderboard. The form judge has no lawyer-labelled accuracy measurement.
  • Uniform document reading: Every model uses the same document-reading system, including for scanned pages. This keeps the focus on legal work but does not test a model or provider’s own document-vision capability.
  • Standardized setup: Instructions, tools, time limits, and tool-call budgets are held constant. Models that benefit from different prompts, tools, or settings may perform differently in practice.
  • Limited consistency measurement: Each task is attempted twice, so the scores capture whether a model repeats its success across two independent attempts, but not wider run-to-run variation.
  • Cost estimates: Cost is priced per workflow from the retained token counts of each attempt at standard API rates, then averaged over the attempts with complete token records. Grading costs are excluded. Actual billing and enterprise discounts may differ.

Get benchmarked for the next quarterly awards