Evaluation Methodology & Disclosure
Purpose and independence
Legal Benchmarks evaluates AI applications on real-world legal tasks. Every application receives the same tasks under the same conditions, is graded against the same lawyer-authored criteria and is run twice, independently.
Applications evaluated
The evaluation covers three types of participants:
- Legal AI applications
- General-purpose AI applications
- AI-native law firms (human + AI)
In a vendor's private report, competing legal AI applications are represented only by the aggregate Legal AI average. No legal AI application's individual result is disclosed to another vendor.
Task scope
The task set comprises demanding, real-world legal work contributed by practitioners in the Legal Benchmarks community. It covers contract drafting and information extraction across 21 jurisdictions, including the United States, the United Kingdom, India, Singapore, the Netherlands and Oman.
Documents are presented as they appear in practice, including native Word files, PDFs, scans, images and documents containing tracked changes. They are not pre-converted into plain text. An application's ability to read and work with the original file is therefore part of the capability being evaluated.
Contract drafting task types
| Task type | Description | Example |
|---|---|---|
| Basic clause drafting | Generating standard clauses that follow widely accepted market practice. | Draft a confidentiality clause requiring both parties to keep information secret for 3 years. |
| Template-based drafting | Adapting an existing contract template based on the provided facts. | Adapt a services agreement template based on the party details. |
| Bespoke drafting | Drafting bespoke clauses or agreements based on specific commercial arrangements. | Draft a revenue-sharing clause for a joint venture where Party A contributes technology and Party B provides distribution, with profits split 70/30. |
Data extraction task types
| Task type | Description | Example |
|---|---|---|
| Single-document extraction | Locating and reporting specific terms or data points from one document. | Identify the governing law and venue provisions in a services agreement. |
| Multi-document extraction | Consolidating information spread across a set of related documents. | Extract the transaction details recorded across a set of related property documents. |
| Imperfect-source extraction | Extracting from documents as they exist in practice: scans, images and files with tracked changes. | Extract the key dates and payment terms from a scanned, partially legible agreement. |
The High Bar
A task enters The High Bar when no more than 2 of the 10 evaluated applications completed it in either run. Applied to the 61 evaluated tasks, this yields 34 High Bar tasks; each application is scored over 68 attempts (34 tasks × 2 runs).
The High Bar metric is the attempt-based Pass rate over those 68 attempts.
Every task has a fixed set of binary pass/fail criteria authored by lawyers. These criteria define what a high-quality answer must include and must avoid. The criteria recognise that legal work does not always have a single acceptable answer. Where several approaches would be professionally defensible, the criteria accommodate those approaches while identifying the specific omissions and errors that should fail.
Each task undergoes legal review before evaluation begins, including verification of its source materials, answer key and pass/fail criteria.
Output collection
AI applications are tested through their normal product interfaces, with documents uploaded in their native formats. The application is configured and used as an ordinary customer would use it unless a different configuration is expressly disclosed.
The complete task set is run twice, and each run is collected independently. The score is calculated as the average of the 2 runs.
Evaluation and grading
Output quality is reported on two separate axes:
- Substance: Is the work substantively correct and complete?
- Form: Could a legal practitioner use the work product as delivered?
The two axes are not combined into a single score. An output can be polished but legally unreliable, or substantively strong but verbose and formatted incorrectly, leading to substantial editing. Reporting them separately preserves that distinction.
Substance
Every output is graded against the fixed pass/fail criteria for the task. Each applicable criterion receives a pass or fail verdict.
Two substance measures are reported:
- Criteria met: The percentage of applicable lawyer-authored criteria satisfied.
- Pass rate: The percentage of tasks for which every applicable criterion was satisfied.
A task passes only when all applicable criteria are met.
Form
Every output is assessed using task-agnostic 1-to-3 rubrics, where 3 is the best score and 1 the worst, for:
- Clarity
- Length
- Structure and formatting
Cross-judge validation
Substance grading is performed independently by two LLM judges: GPT-5.6 Sol at high reasoning effort and Claude Opus 4.8.
Form grading is performed independently by GPT-5.6 Sol and Claude Sonnet 4.6.
For substance, both judges assess the output against the same fixed binary criteria; they are not asked to express a preference or reward writing style. For form, both judges use the same task-agnostic rubrics.
Human adjudication
For substance, all judge disagreements are logged, and any disagreement that could change whether a task passes is escalated to a qualified lawyer for a final, binding determination. The lawyer reviews the output blind, without knowing which application produced it or which judge issued each verdict; the human determination takes precedence and is preserved for audit.
For form, a two-point disagreement is escalated to a qualified lawyer for final adjudication.
Results and disclosure
Comparison groups
- Legal AI average: the other evaluated legal AI applications, excluding the vendor receiving the private report.
- General-purpose average: the evaluated general-purpose applications.
- Industry average: all evaluated applications in the cohort.
What's public: Named results for general-purpose AI applications and open-source references, the aggregate Legal AI average, published certifications and badges, and the evaluation period and context needed to interpret them.
A legal AI application's own name and result become public only where the vendor opts in to be public or accepts a publicly awarded certification. Legal Benchmarks may publish anonymised excerpts of application outputs to illustrate findings and methodology, provided the excerpt cannot be traced to the application that produced it.
What's confidential: The vendor's private report, per-task results and grading records, failure-mode diagnostics and improvement analysis. Confidential information is published only with the vendor's authorisation, and a private report never identifies another legal AI application or discloses its individual result; competitors appear only as the Legal AI average.
Retesting
A vendor may request one retest per quarter, subject to the quarter's publication deadline and to Legal Benchmarks' discretion; additional fees may apply. The retest runs the complete task set again under the same conditions, and the new result replaces the earlier one in the quarter's published results.
A retest after the deadline enters the following quarter's evaluation.
Certifications
Certifications are awarded quarterly by category. An application may earn multiple certifications.
| Certification | What it recognises |
|---|---|
| Best Legal AI Application (Overall)Industry Leader | The application meets the highest certification threshold across the evaluation as a whole, considering both substance and form across all tested categories. |
| Contract DraftingIndustry Leader | The application meets the highest threshold for producing and revising legal drafting, including clauses, amendments, redlines and complete agreements. |
| Information ExtractionIndustry Leader | The application meets the highest threshold for accurately locating, interpreting and reporting information from legal documents, including multi-document and imperfect-file tasks. |
| The High BarContender | The application meets the highest threshold on the benchmark's most difficult and differentiating tasks. This subset is designed to distinguish exceptional performance from competence on standard tasks. |
| User ExperienceIndustry Leader | The application meets the highest threshold for the practical experience of completing legal work through the product, including usability, document handling and the quality of the delivered work product. |
If no application reaches the required threshold in a category, no certification is awarded for that category that quarter.
During an active quarter, participating applications may be designated Contenders. Certifications are confirmed when the quarter closes.