Skip to content

Methodology

How Legal Benchmarks evaluates AI on real-world legal work: what we test, how we run models and applications, how we score their output, and how we validate the graders behind both leaderboards.
Framework · August 2026

What we benchmark

Legal Benchmarks publishes two leaderboards built on one task set.

  • The Models leaderboard tests AI models directly, through a controlled evaluation harness.
  • The Applications leaderboard tests complete AI products through their normal user interfaces, exactly as a customer would use them. It covers legal AI applications, general-purpose AI applications, and AI-native law firms (human + AI).

Both leaderboards use the same tasks, the same lawyer-authored criteria, and the same scoring axes, so a model result and an application result are measuring the same legal work.

A model is tested on its raw capability under identical conditions; an application is tested as a product, including its own document handling, configuration, and interface.

Task scope

The task set comprises demanding, real-world legal work contributed by practitioners in the Legal Benchmarks community. It spans two categories, with tasks contributed from 21 jurisdictions, including the United States, the United Kingdom, India, Singapore, the Netherlands and Oman, with a US and UK concentration.

Contract workflows cover producing or amending contract language from a lawyer’s instruction.

Task typeDescriptionExample
Basic clause draftingGenerating standard clauses that follow widely accepted market practice.Draft a confidentiality clause requiring both parties to keep information secret for 3 years.
Template-based draftingAdapting an existing contract template based on the provided facts.Adapt a services agreement template based on the party details.
Bespoke draftingDrafting bespoke clauses or agreements based on specific commercial arrangements.Draft a revenue-sharing clause for a joint venture where Party A contributes technology and Party B provides distribution, with profits split 70/30.

Data extraction covers answering questions about a document, such as locating clauses, defined terms, governing law, values, and obligations.

Task typeDescriptionExample
Single-document extractionLocating and reporting specific terms or data points from one document.Identify the governing law and venue provisions in a services agreement.
Multi-document extractionConsolidating information spread across a set of related documents.Extract the transaction details recorded across a set of related property documents.
Imperfect-source extractionExtracting from documents as they exist in practice: scans, images and files with tracked changes.Extract the key dates and payment terms from a scanned, partially legible agreement.

Documents are presented as they appear in practice, including native Word files, PDFs, scans, images and documents containing tracked changes. They are not pre-converted into plain text.

Where the tasks come from

  • Practitioner-authored tasks make up the core of the set. Lawyers from the global Legal Benchmarks community wrote the instructions, reference files, and answer keys based on real scenarios from their own work.
  • Synthetic tasks make up the rest. We generate them from the failure modes identified in our earlier Data Extraction and Contract Workflows reports, and from published legal-AI benchmarking work in the industry. Every synthetic task is validated by a human expert before it enters the set.

Every task has a fixed set of binary pass/fail criteria authored by lawyers. These criteria define what a high-quality answer must include and must avoid, and they recognise that legal work does not always have a single acceptable answer: where several approaches would be professionally defensible, the criteria accommodate those approaches while identifying the specific omissions and errors that should fail. Each task undergoes legal review before evaluation begins, including verification of its source materials, answer key and pass/fail criteria.

This is the third installment in a series we have been publishing since April 2025. Our Phase 1 report focused on data extraction, the Phase 2 report expanded to contract workflows, a human-lawyer baseline, and the two-axis split between substance and form that this benchmark inherits. The current task set consolidates both into a single, repeatable evaluation.

How we run the evaluation

Every subject on either leaderboard is evaluated by us, under our control. Nobody self-reports results.

Models

Each task runs in a sealed workspace. The model receives the lawyer’s instruction and the source documents, works alone with a small, fixed set of file tools, and submits its answer together with any documents it produced. The workspace has no web access, every model gets an identical setup, and every attempt starts fresh with no memory of any other task. Runs operate under fixed time and tool-call budgets; a model that cannot finish the work inside them has failed the task.

The harness holds everything constant except the model. Where a provider exposes a reasoning-effort control we request the high setting; otherwise the provider’s defaults apply.

Applications

Applications are tested through their normal product interfaces, with documents uploaded in their native formats. The application is configured and used as an ordinary customer would use it unless a different configuration is expressly disclosed. An application’s ability to read and work with the original file is part of the capability being evaluated.

The complete task set is run twice, and each run is collected independently. The score is the average of the two runs.

Alongside substance and form, applications are also assessed on product experience: the practical experience of completing legal work through the product, including usability and document handling.

The two evaluations at a glance

ModelsApplications
SubjectThe model itselfThe complete product
AccessSealed harness workspace with fixed file toolsThe product’s own interface, as a customer uses it
Document handlingUniform reading tool, identical for every modelThe application’s own ingestion and file handling
AttemptsOne per taskTwo per task, averaged
What it measuresRaw model capability under identical conditionsEnd-to-end product performance, including document handling and interface

Measuring output quality

Quality is reported on two independent axes, substance and form, kept separate rather than blended into a single score. An output can be polished but legally unreliable, or substantively strong but verbose and poorly formatted. Reporting the axes separately preserves that distinction.

Substance asks whether the legal work is correct, complete, and responsive. It also captures honesty about what the AI does not know rather than inventing. Two substance measures are reported:

  • Criteria met: the percentage of applicable lawyer-authored criteria satisfied.
  • Task pass rate: the percentage of tasks for which every applicable criterion was satisfied. A task passes only when all applicable criteria are met.

Form asks whether the output is ready to use as delivered. Every output is assessed on task-agnostic 1-to-3 rubrics for clarity, length, and structure and formatting, where 1 means needs rework before use, 2 means usable with edits, and 3 means use as-is.

Grading and validation

Substance is graded against each task’s fixed binary criteria by cross-family LLM judges; the judges apply the criteria as written and are not asked to express a preference or reward writing style. Form is scored by a cross-family panel using the same task-agnostic rubrics, with panel scores averaged into a consensus. Judges grade the submission, meaning the final answer and any files delivered, blind to the reasoning that produced it, and the two axes are graded independently: neither judge sees the other’s scores or comments.

  • Cross-judge agreement. Judge disagreements are logged. Agreement is highest on clarity and structure and lowest on length, the most subjective sub-criterion and the one most worth a human eye.
  • Human adjudication. For substance, any disagreement that could change whether a task passes is escalated to a qualified lawyer for a final, binding determination. The lawyer reviews the output blind, without knowing which subject produced it or which judge issued each verdict; the human determination takes precedence and is preserved for audit. For form, a maximum-spread disagreement is escalated the same way.
  • Human calibration. Practitioners spot-check judge verdicts on a sample of every run, and confirmed corrections carry into the published scores.
  • Favoritism. We monitor whether any grader scores its own maker’s models higher, and we publish per-sub-criterion scores so readers can check the pattern themselves.

Results and rankings

Contract workflows and data extraction are reported per category on both leaderboards, alongside overall figures, with clarity, length, and structure also reported individually because they fail differently. Rankings are refreshed each time we re-run the benchmark.

Publication schedule (applications only)

Certifications are awarded quarterly on performance alone; the first awards are issued in September 2026. During an active quarter, participating legal AI applications are evaluated as contenders, and their results are confirmed when the quarter closes. The leaderboard today carries named results for general-purpose applications and open-source references; legal AI application results follow the disclosure rules below.

Disclosure (applications only)

What’s public: named results for general-purpose AI applications and open-source references, the aggregate Legal AI average, published certifications and badges, and the evaluation period and context needed to interpret them. A legal AI application’s own name and result become public only where the vendor opts in or accepts a publicly awarded certification. Legal Benchmarks may publish anonymised excerpts of outputs to illustrate findings, provided the excerpt cannot be traced to the application that produced it.

What’s confidential: the vendor’s private report, per-task results and grading records, failure-mode diagnostics and improvement analysis. A private report never identifies another legal AI application or discloses its individual result; competitors appear only as the Legal AI average.

Certification categories, thresholds, and retesting terms are set out on Get benchmarked.

Accessing the tasks

The task set, reference documents, and per-task criteria are not publicly distributed. Releasing them in full would let vendors train or tune against the benchmark, which would quickly erode the signal we are trying to publish. Researchers and academics who want to inspect the methodology in more detail can request access by getting in touch.

Vendors who want their application measured against the set can submit it through Get benchmarked, and we run the evaluation on their behalf.

Limitations

The benchmark gives a useful read on where models and applications stand today, but it has limits worth flagging.

  • Coverage. The set leans commercial, IP, employment, M&A, and competition work, and it is English only with a US and UK skew. Practice areas and jurisdictions outside that scope are thin or absent.
  • Single assignment only. Every task is one instruction and one submission. We do not yet test back-and-forth refinement with a lawyer, multi-turn workflows, or longer matters, which is closer to how lawyers actually use these tools.
  • Snapshot, not a trend. Models and applications change month to month. The leaderboards reflect a point-in-time run, and a result from one month may not hold the next.
  • LLM judges. Grading is performed by LLM judges validated by human experts on a spot-check and escalation basis, not on every task.
  • Uniform document reading (models only). Every model reads source files through the same reading tool, including scanned pages. That levels the playing field on document intake and keeps the score focused on the legal work, but it means the Models leaderboard does not measure a model’s own document-vision capability, and scan-heavy tasks are easier than they would be if models handled scans themselves. We report results split by whether a task’s sources include a scanned page. The Applications leaderboard does test native document handling.
  • Standardized setup (models only). Instructions, workspace, and run settings are identical across models, so the only variable is the model itself. A model that benefits from a more tailored prompt or different settings may look weaker here than it does in practice.
  • Cost figures (models only). Per-task cost is priced at listed API rates at the time of the run so figures are comparable across models. Actual billed cost can differ, and enterprise discounts are not reflected.
  • Consistency (models only). Models are graded on a single attempt per task; applications on two. We do not yet measure output variance more broadly, which matters for anyone relying on these tools day to day.

Changelog

August 2026 (Applications)

  • Launched the Applications leaderboard, with named results for general-purpose applications and open-source references. Legal AI applications and AI-native law firms are being evaluated in the current cycle; certifications and opted-in named results are issued in September 2026.

August 2026 (Models)

  • Moved model evaluation to an agentic harness: models now work in a sealed workspace with a uniform tool surface, including a shared document-reading tool with OCR, and submit file deliverables instead of answering in a single API call. Substance is now reported as both criteria met and a strict task pass rate, with overall figures alongside the per-category rankings. Refreshed the task set and expanded the roster. Results from this cycle are not comparable with earlier cycles.

July 2026 (Models)

  • Added Claude Opus 5, Kimi K3, GPT-5.6 Sol, and Grok 4.5 to the board.

June 2026 (Models)

  • Expanded the ranked model field: added Claude Fable 5, Gemini 3.5 Flash, GPT-5.4-mini, DeepSeek V4 Pro, Qwen3.7 Max, and Claude Opus 4.8; retired Claude Opus 4.7, Gemini 3 Flash, and GPT-4o-mini. Switched document handling so models receive documents rather than extracted text. Added the Limitations section.