Introduction
Designed and validated by 100+ legal and technology leaders from the buy side across the globe.
Legal teams have well-established playbooks for hiring humans, but none for hiring AI agents. This is that playbook, built by the legal community.
Our Principles
Independent
Lawyer-led
Grounded in real work
Why Legal AI Procurement is Broken
- 88%*
- Of legal teams are not committed to their current AI vendor. Many legal teams are still searching for tools that meet their needs.
- 63%*
- Of legal teams evaluate AI tools without IT or security involvement. Many procurement decisions happen without a structured technical review.
- 800+**
- AI tools now target legal teams. Many promise similar capabilities, making meaningful vendor comparison difficult.
The result: months of duplicated effort, inconsistent evaluation, and decisions driven by demos and marketing rather than evidence.
The Legal AI Evaluation Framework
The legal community's first open-access evaluation framework, built to help legal teams make defensible, evidence-based procurement decisions.
- 100+
- Legal leaders across 82 organizations and law firms
- 25
- Countries represented
- 100
- Evaluation sub-tests across 8 core criteria
The 8 Core Evaluation Criteria
| # | Criterion | What it covers |
|---|---|---|
| 01 | Strategic Fit | Alignment to legal use cases, IT systems, jurisdictions, languages, and long-term legal team needs. |
| 02 | Functionality | User interface design, workflow integration, customisation, input handling, and real-world usability within existing technology stacks and processes. |
| 03 | Robustness | Factual accuracy, completeness, instruction fidelity, verifiability, citation quality, and consistency of AI-generated outputs across legal tasks. |
| 04 | Security | Architecture and data flow transparency, access control, retrieval boundaries, adversarial resistance, AI safety, and alignment with the organisation’s security policy requirements. |
| 05 | Data Privacy | Data use restrictions (including no-training provisions covering aggregated and derived data), deletion, localisation, vector embedding governance, and sub-processor practices. |
| 06 | Vendor Risk | Licensing, contractual security commitments, data portability, audit trails, transparency, vendor track record, business continuity, and incident response. |
| 07 | Adoption Support | Training, support, issue resolution, documentation, and usage reporting to support rollout and sustained adoption. |
| 08 | Cost & Resourcing | Pricing model, total cost of ownership, and internal operational capacity required to deploy and maintain the tool. |
The 3 Stages of Legal AI Evaluation
Like hiring a knowledge worker, each stage demands a different level of scrutiny. The toolkit turns the framework into practical tools for buyers at every decision gate.
Stage 1: Pre-demo research
Pre-Demo Checklist. Pass/fail screening to decide if a demo is worth booking.
Hiring analogy: Resume screening. Decision gate: Proceed / Do Not Proceed.
Stage 2: Live demo
Demo Scorecard. Scored validation of whether the tool performs as claimed.
Hiring analogy: Interview. Decision gate: Proceed / Do Not Proceed.
Stage 3: Hands-on pilot
Pilot Scorecard. Evidence-based, weighted evaluation using real workflows.
Hiring analogy: Working trial. Decision gate: Proceed / Do Not Proceed.
How to Use This Toolkit
Start at the stage that matches where you are in the buying process.
This toolkit applies whether you are:
- Starting a new search for a legal AI tool
- Midway through vendor conversations and need structure
- Running a pilot and need a consistent way to compare tools
- Reassessing or renewing a contract with your current provider
Researching tools?
Sitting in demos?
Trialing a tool?
Starting from scratch?
Then make it yours
How We Built This
This framework was built from the ground up by the people who do this work. It is a living framework that will evolve as the tools, models, and market change.
November to December 2025 · done
Framework Development
February 21, 2026 · done
Community Feedback
Late February 2026 · done
v1 Finalisation
March 11, 2026 · current
Publication and Launch
From March 2026 · ongoing
What Comes Next
Built by Practitioners Across Legal, IT, Security, and Privacy
The framework is built and maintained by legal professionals, technologists, security experts, and privacy specialists. The core team is responsible for day-to-day development, content, publication, and coordination across 120+ contributors.
Core team
The core team designed and published the framework and handles day-to-day development, content, and coordination.
Anna Guo
Founder, Legal Benchmarks
Elgar Weijtmans
Technologist & Former Lawyer
Roel Schrijvers
General Counsel
Sunny Kim
Comms Lead