Skip to content

Research / DELTA

Introducing DELTA

A public benchmark for evaluating whether AI can do real Dutch legal work with accuracy, judgment, and legal taste.

DELTA is a public, practitioner-led benchmark measuring how foundation models perform on real Dutch legal work. It asks not only whether AI gets the law right, but whether its work demonstrates legal taste: the professional judgment to identify what matters and deliver an answer a lawyer could use.

Named for the Dutch delta and the symbol for change, DELTA measures progress against a standard rooted in Dutch legal practice.

Its first public release includes an open legal-research task set and a survey of the Dutch legal community: 115 legal professionals on how they use AI, and what they expect from it.

AI will play an increasingly important role in legal research, but lawyers must remain critical about the quality, reliability and verifiability of its output. Independent evaluation helps us understand not only what legal AI can do, but also where its limits lie.
Jeroen MaasPartner, DM Advocaten, Belastingadviseurs, Mediators

What legal taste means

Legal taste is not personal style. It is the observable professional judgment involved in identifying the decisive issues, separating material points from distractions, handling uncertainty, calibrating caveats and delivering a usable answer.

A legal answer can be fluent and well presented while still reaching the wrong conclusion, overlooking a material issue or failing to cite the authority needed to support it. DELTA assesses each task in three categories. Legal taste crosses substance and form, while citation anchors the work in legal authority.

Substance

Whether the analysis reaches legally correct and professionally defensible conclusions, covers the material rules and issues and addresses the assignment completely.

Citation

Whether the answer identifies the necessary authorities accurately and uses them to support the relevant legal propositions.

Form

Whether the answer is clearly structured, appropriately prioritised, proportionate, properly qualified and usable by a practitioner.

Grounded in Dutch legal practice

The first public tasks were selected from a growing collection of more than 200 legal-research assignments reflecting questions encountered in Dutch practice. Practising Dutch lawyers reviewed the questions and assessment criteria and scored model outputs, informing the development and calibration of the framework.

Across the wider DELTA project, 12 partner law firms and 115 legal professionals contributed through task review, model evaluation and adoption research. The survey broadened that evidence across the Dutch legal community, examining how legal professionals use AI, where it fails and where human judgment must remain decisive.

  • With so many vendors and approaches, independent evaluation matters at the organisational level, not just tool by tool. Proper grounding and clearer information on accuracy, precision and failure modes are what make AI something an organisation can rely on, not just something that sounds convincing.
    Mark ZijlstraHead of Legal Technology, ICTRecht

Founding Research Partner

Zeno is a Dutch legal AI company building legal intelligence for professional legal work. Its technology models how legislation, case law and doctrine relate to one another, supporting legal research and analysis with traceable reasoning and sources while keeping the lawyer in control.

As DELTA's Founding Research Partner, Zeno contributed technical expertise in legal AI systems, evaluation design and benchmark development.

Our research collaboration with Legal Benchmarks was grounded in a simple principle: progress in legal AI should be measured against the same expectations lawyers bring to real work. DELTA gives the legal and AI communities a transparent, shared way to assess current capabilities, track improvement and focus development where it matters most.
Nordin BouchritCPO at Zeno

Explore DELTA

Available now

Legal Research Task Set

Open-ended research questions drawn from Dutch legal practice, assessed against 273 criteria for substance, citation and form across 9 areas of Dutch law.

Explore the task set on GitHub
Available now

Dutch Legal AI Adoption Survey

Companion research examining how Dutch legal professionals use AI for legal research, which failures they consider most serious and what evidence they want when assessing legal AI.

Read the survey report
Coming soon

DELTA Dutch Legal Research Benchmark Report

An independent evaluation of leading foundation models operating through the same legal-research harness, covering complete-task performance, legal substance, citation, professional judgment, form, cost and time.

Join the research community

An open benchmark becomes stronger when the community can examine its assumptions, challenge its criteria and extend its coverage. Review the tasks, examine the criteria or contribute a perspective on what legal taste requires in practice.