Skip to content

The Evaluation Framework

The legal community's playbook for evaluating AI agents with the same rigor legal teams apply when hiring people.
Framework · Feb 2026

Introduction

Designed and validated by 100+ legal and technology leaders from the buy side across the globe.

Legal teams have well-established playbooks for hiring humans, but none for hiring AI agents. This is that playbook, built by the legal community.

Our Principles

Independent

No vendor sponsorship or commercial influence.

Lawyer-led

Designed, run, and reviewed by practicing legal professionals, with 500+ involved to date.

Grounded in real work

Benchmarks built from actual legal tasks, not synthetic test questions.

Why Legal AI Procurement is Broken

88%*
Of legal teams are not committed to their current AI vendor. Many legal teams are still searching for tools that meet their needs.
63%*
Of legal teams evaluate AI tools without IT or security involvement. Many procurement decisions happen without a structured technical review.
800+**
AI tools now target legal teams. Many promise similar capabilities, making meaningful vendor comparison difficult.

The result: months of duplicated effort, inconsistent evaluation, and decisions driven by demos and marketing rather than evidence.

The Legal AI Evaluation Framework

The legal community's first open-access evaluation framework, built to help legal teams make defensible, evidence-based procurement decisions.

100+
Legal leaders across 82 organizations and law firms
25
Countries represented
100
Evaluation sub-tests across 8 core criteria

The 8 Core Evaluation Criteria

The eight criteria used to evaluate legal AI tools
#CriterionWhat it covers
01Strategic FitAlignment to legal use cases, IT systems, jurisdictions, languages, and long-term legal team needs.
02FunctionalityUser interface design, workflow integration, customisation, input handling, and real-world usability within existing technology stacks and processes.
03RobustnessFactual accuracy, completeness, instruction fidelity, verifiability, citation quality, and consistency of AI-generated outputs across legal tasks.
04SecurityArchitecture and data flow transparency, access control, retrieval boundaries, adversarial resistance, AI safety, and alignment with the organisation’s security policy requirements.
05Data PrivacyData use restrictions (including no-training provisions covering aggregated and derived data), deletion, localisation, vector embedding governance, and sub-processor practices.
06Vendor RiskLicensing, contractual security commitments, data portability, audit trails, transparency, vendor track record, business continuity, and incident response.
07Adoption SupportTraining, support, issue resolution, documentation, and usage reporting to support rollout and sustained adoption.
08Cost & ResourcingPricing model, total cost of ownership, and internal operational capacity required to deploy and maintain the tool.

The 3 Stages of Legal AI Evaluation

Like hiring a knowledge worker, each stage demands a different level of scrutiny. The toolkit turns the framework into practical tools for buyers at every decision gate.

Stage 1: Pre-demo research

Pre-Demo Checklist. Pass/fail screening to decide if a demo is worth booking.

Hiring analogy: Resume screening. Decision gate: Proceed / Do Not Proceed.

Get the full toolkit

Stage 2: Live demo

Demo Scorecard. Scored validation of whether the tool performs as claimed.

Hiring analogy: Interview. Decision gate: Proceed / Do Not Proceed.

Get the full toolkit

Stage 3: Hands-on pilot

Pilot Scorecard. Evidence-based, weighted evaluation using real workflows.

Hiring analogy: Working trial. Decision gate: Proceed / Do Not Proceed.

Get the full toolkit

How to Use This Toolkit

Start at the stage that matches where you are in the buying process.

This toolkit applies whether you are:

  • Starting a new search for a legal AI tool
  • Midway through vendor conversations and need structure
  • Running a pilot and need a consistent way to compare tools
  • Reassessing or renewing a contract with your current provider

Researching tools?

Start with the Pre-Demo Checklist (Stage 1)

Sitting in demos?

Use the Demo Scorecard (Stage 2)

Trialing a tool?

Use the Pilot Scorecard (Stage 3)

Starting from scratch?

Work through all three stages in order.

Then make it yours

Adapt the criteria to your team's priorities. The toolkits are starting points, not rigid templates: the framework gives you structure, and you decide what matters most.

How We Built This

This framework was built from the ground up by the people who do this work. It is a living framework that will evolve as the tools, models, and market change.

November to December 2025 · done

Framework Development

Synthesised legal AI evaluation approaches used in real vendor selection processes and assembled the first draft of the framework.

February 21, 2026 · done

Community Feedback

Opened the draft for community input. More than 100 legal leaders contributed feedback.

Late February 2026 · done

v1 Finalisation

Incorporated community feedback into Version 1 of the Legal AI Evaluation Framework and the accompanying Evaluation ToolKit.

March 11, 2026 · current

Publication and Launch

Public release of the Legal AI Evaluation Framework v1, together with the practical toolkits that operationalise the framework for real evaluation workflows.

From March 2026 · ongoing

What Comes Next

Gathering post-launch feedback and releasing additional practical resources, including a vendor questionnaire, the evaluation framework report, a quick-reference cheat sheet, and community insights.

Built by Practitioners Across Legal, IT, Security, and Privacy

The framework is built and maintained by legal professionals, technologists, security experts, and privacy specialists. The core team is responsible for day-to-day development, content, publication, and coordination across 120+ contributors.

Core team

The core team designed and published the framework and handles day-to-day development, content, and coordination.

Anna Guo

Anna Guo

Founder, Legal Benchmarks

Over the past 13 months, Anna has benchmarked AI performance in real legal workflows alongside 500+ legal professionals. That work has shaped hundreds of procurement decisions around AI tooling.
Elgar Weijtmans

Elgar Weijtmans

Technologist & Former Lawyer

Led legal AI procurement and evaluated 40+ tools last year. Brings an end-to-end view of AI assessment, from screening and testing to piloting and making the final tool decision.
Roel Schrijvers

Roel Schrijvers

General Counsel

Focused on data security and operational risk in AI systems. Pushes the team to ask the questions most organisations miss around security, safety, and deployment risk.
Sunny Kim

Sunny Kim

Comms Lead

Leads communications and PR for the framework. Brings experience in legal and professional services communications to help the community understand and engage with the project.

Frequently Asked Questions

Put the framework to work