Legal Benchmarks
Legal Benchmarks · Application Evaluation
legalbenchmarks.ai

Evaluation Methodology & Disclosure

Section 01

Purpose and independence

Legal Benchmarks evaluates AI applications on real-world legal tasks. Every application receives the same tasks under the same conditions, is graded against the same lawyer-authored criteria and is run twice, independently.

Section 02

Applications evaluated

The evaluation covers three types of participants:

In a vendor's private report, competing legal AI applications are represented only by the aggregate Legal AI average. No legal AI application's individual result is disclosed to another vendor.

Section 03

Task scope

The task set comprises demanding, real-world legal work contributed by practitioners in the Legal Benchmarks community. It covers contract drafting and information extraction across 21 jurisdictions, including the United States, the United Kingdom, India, Singapore, the Netherlands and Oman.

Documents are presented as they appear in practice, including native Word files, PDFs, scans, images and documents containing tracked changes. They are not pre-converted into plain text. An application's ability to read and work with the original file is therefore part of the capability being evaluated.

Contract drafting task types

Task typeDescriptionExample
Basic clause draftingGenerating standard clauses that follow widely accepted market practice.Draft a confidentiality clause requiring both parties to keep information secret for 3 years.
Template-based draftingAdapting an existing contract template based on the provided facts.Adapt a services agreement template based on the party details.
Bespoke draftingDrafting bespoke clauses or agreements based on specific commercial arrangements.Draft a revenue-sharing clause for a joint venture where Party A contributes technology and Party B provides distribution, with profits split 70/30.

Data extraction task types

Task typeDescriptionExample
Single-document extractionLocating and reporting specific terms or data points from one document.Identify the governing law and venue provisions in a services agreement.
Multi-document extractionConsolidating information spread across a set of related documents.Extract the transaction details recorded across a set of related property documents.
Imperfect-source extractionExtracting from documents as they exist in practice: scans, images and files with tracked changes.Extract the key dates and payment terms from a scanned, partially legible agreement.

The High Bar

A task enters The High Bar when no more than 2 of the 10 evaluated applications completed it in either run. Applied to the 61 evaluated tasks, this yields 34 High Bar tasks; each application is scored over 68 attempts (34 tasks × 2 runs).

The High Bar metric is the attempt-based Pass rate over those 68 attempts.

Every task has a fixed set of binary pass/fail criteria authored by lawyers. These criteria define what a high-quality answer must include and must avoid. The criteria recognise that legal work does not always have a single acceptable answer. Where several approaches would be professionally defensible, the criteria accommodate those approaches while identifying the specific omissions and errors that should fail.

Each task undergoes legal review before evaluation begins, including verification of its source materials, answer key and pass/fail criteria.

Section 04

Output collection

AI applications are tested through their normal product interfaces, with documents uploaded in their native formats. The application is configured and used as an ordinary customer would use it unless a different configuration is expressly disclosed.

The complete task set is run twice, and each run is collected independently. The score is calculated as the average of the 2 runs.

Section 05

Evaluation and grading

Output quality is reported on two separate axes:

The two axes are not combined into a single score. An output can be polished but legally unreliable, or substantively strong but verbose and formatted incorrectly, leading to substantial editing. Reporting them separately preserves that distinction.

Substance

Every output is graded against the fixed pass/fail criteria for the task. Each applicable criterion receives a pass or fail verdict.

Two substance measures are reported:

A task passes only when all applicable criteria are met.

Form

Every output is assessed using task-agnostic 1-to-3 rubrics, where 3 is the best score and 1 the worst, for:

Cross-judge validation

Substance grading is performed independently by two LLM judges: GPT-5.6 Sol at high reasoning effort and Claude Opus 4.8.

Form grading is performed independently by GPT-5.6 Sol and Claude Sonnet 4.6.

For substance, both judges assess the output against the same fixed binary criteria; they are not asked to express a preference or reward writing style. For form, both judges use the same task-agnostic rubrics.

Human adjudication

For substance, all judge disagreements are logged, and any disagreement that could change whether a task passes is escalated to a qualified lawyer for a final, binding determination. The lawyer reviews the output blind, without knowing which application produced it or which judge issued each verdict; the human determination takes precedence and is preserved for audit.

For form, a two-point disagreement is escalated to a qualified lawyer for final adjudication.

Section 06

Results and disclosure

Comparison groups

What's public: Named results for general-purpose AI applications and open-source references, the aggregate Legal AI average, published certifications and badges, and the evaluation period and context needed to interpret them.

A legal AI application's own name and result become public only where the vendor opts in to be public or accepts a publicly awarded certification. Legal Benchmarks may publish anonymised excerpts of application outputs to illustrate findings and methodology, provided the excerpt cannot be traced to the application that produced it.

What's confidential: The vendor's private report, per-task results and grading records, failure-mode diagnostics and improvement analysis. Confidential information is published only with the vendor's authorisation, and a private report never identifies another legal AI application or discloses its individual result; competitors appear only as the Legal AI average.

Section 07

Retesting

A vendor may request one retest per quarter, subject to the quarter's publication deadline and to Legal Benchmarks' discretion; additional fees may apply. The retest runs the complete task set again under the same conditions, and the new result replaces the earlier one in the quarter's published results.

A retest after the deadline enters the following quarter's evaluation.

Section 08

Certifications

Certifications are awarded quarterly by category. An application may earn multiple certifications.

CertificationWhat it recognises
Best Legal AI Application (Overall)Industry LeaderThe application meets the highest certification threshold across the evaluation as a whole, considering both substance and form across all tested categories.
Contract DraftingIndustry LeaderThe application meets the highest threshold for producing and revising legal drafting, including clauses, amendments, redlines and complete agreements.
Information ExtractionIndustry LeaderThe application meets the highest threshold for accurately locating, interpreting and reporting information from legal documents, including multi-document and imperfect-file tasks.
The High BarContenderThe application meets the highest threshold on the benchmark's most difficult and differentiating tasks. This subset is designed to distinguish exceptional performance from competence on standard tasks.
User ExperienceIndustry LeaderThe application meets the highest threshold for the practical experience of completing legal work through the product, including usability, document handling and the quality of the delivered work product.

If no application reaches the required threshold in a category, no certification is awarded for that category that quarter.

During an active quarter, participating applications may be designated Contenders. Certifications are confirmed when the quarter closes.