Skip to content

In Brazilian Law, the Right Answer Has an Expiry Date

The most legible jurisdiction in the world for evaluating legal AI is also the only major one without an instrument of its own.
Marcus Camargo

Marcus Camargo

Published August 28, 2026

In Brazilian Law, the Right Answer Has an Expiry Date

Task sets assembled for common law systems and for transactional practice do not transfer to Brazil. The economic centre here is mass litigation; binding theses have their temporal effects modulated; and forum divergence has no superior court positioned to resolve it across a substantial share of the volume.

What has gone unnoticed is that these same difficulties coexist with evaluation conditions no other major jurisdiction offers. Many core Brazilian legal identifiers are numbered and partly machine-verifiable. The standard for adequate judicial reasoning is codified in statute, enumerated in six items. An official reference corpus is maintained by the judiciary's own training school. And binding regulation of AI in the judiciary has been in force since July 2025.

What is missing is not a methodological condition. It is someone assembling the set, and there is a reason for it to be now.

PART ONE: WHY EXISTING TASK SETS DO NOT TRANSFER

I. THE PREMISE

Latin America is not one market

Benchmarks that report jurisdictional coverage tend to count countries. That is the wrong unit. What determines whether a legal AI system performs in a given place is not the country label but the structure of its legal reasoning: where authority comes from, how stable it is, and what a correct citation looks like.

On that measure Brazil is not adjacent to its neighbours. It is a civil law system with binding precedent grafted onto it, operating in Portuguese, at a case volume that has no parallel anywhere in the world, through an electronic record fragmented across incompatible systems. A task set assembled for Spanish-speaking Latin America does not transfer. Neither does one assembled for Portugal.

The practical consequence is the one that matters commercially: a buyer in São Paulo pays the same licence fee as a buyer in London, for a product whose reliability in Brazilian matters nobody has measured.

II. THE TASK SET

The economic centre is volume, not deals

Published task sets concentrate on transactional and advisory work: M&A diligence, capital markets, funds, private equity, tax memoranda. That reflects where the largest budgets sit, and it is a defensible commercial choice.

It is not, however, where Brazilian legal work happens. The centre of gravity here is high-volume repetitive litigation — consumer, telecommunications, banking, social security, labour — filed and defended at industrial scale by parties who litigate thousands of substantially similar matters at once.

This is not a smaller version of the transactional problem. It is a different one. Transactional work rewards depth on a single complex artefact. Mass litigation rewards consistency across a large population of near-identical ones, where the failure that matters is not a missed nuance but a systematic error replicated ten thousand times before anyone notices.

An agent that drafts an excellent defence once has demonstrated very little. An agent that drafts four hundred without a single misapplied precedent or transposed docket number has demonstrated the thing the market actually buys.

III. GROUND TRUTH

The answer moves after you give it

Brazilian law binds the lower courts through instruments that are themselves in motion. The Supreme Federal Court issues binding summaries and decides questions under general repercussion; the Superior Court of Justice fixes theses under the repetitive appeals procedure; the courts below resolve mass questions through incidents created for that purpose. Each produces a numbered thesis that lower courts are expected to follow.

Two properties of that architecture make evaluation genuinely hard.

First, the theses are revised. A position held for a decade is overturned, and every filing drafted under the old thesis becomes wrong — not badly reasoned, simply wrong.

Second, and more difficult: Brazilian courts routinely modulate the temporal effects of a change. A new thesis may bind prospectively while the old one continues to govern facts that preceded it. The correct answer therefore depends not only on what the law says today, but on the date of the facts, the date of filing, and the date the court chose as its cut-off.

A benchmark that scores against a snapshot of current law will mark correct answers wrong and wrong answers correct, in both directions, without any way to tell which happened.

Ground truth in Brazil is not a value. It is a function of time — and the benchmark has to model it as one.

IV. DIVERGENCE

Divergence with no court to resolve it

Twenty-seven state courts, six federal regional courts, twenty-four labour courts. The same question receives different answers in different regions, and that is ordinary rather than exceptional.

The harder case sits in the small claims courts. Their appeals are decided by panels of first-instance judges rather than by courts of appeal, and a binding summary of the Superior Court of Justice excludes special appeal against those decisions. What remains is constitutional appeal to the Supreme Federal Court, and only where a constitutional question arises. Federal small claims have a uniformisation procedure; the state ones have no national equivalent. The result is that across an enormous share of Brazilian litigation by volume, the interpretation of federal law is simply never unified. Divergence is not a defect awaiting correction. It is the permanent state of the system.

This breaks the assumption underneath most legal benchmarks — that a question has one correct answer against which output can be scored. In Brazil a great many questions have several, each correct in its forum.

V. THE RECORD

Electronic, and therefore fragmented

Brazilian litigation is almost entirely electronic. That sounds like a simplification and is the opposite of one.

The courts did not converge on a single platform. Several systems operate in parallel across the federal, state and labour branches, each exporting the case file with its own conventions for pagination, ordering, signature blocks and the placement of the validation stamp. A practitioner moving between two states reads the same underlying record through two different renderings of it.

Published task sets have begun to include multiple file formats, which is a real advance. But format variety is not the Brazilian problem. The problem is that the same document type arrives structured differently depending on which court produced it, and that scanned material — old filings, powers of attorney, exhibits photographed on a phone — sits inside otherwise clean electronic records without warning.

PART TWO:WHY BRAZIL IS NONETHELESS THE MOST LEGIBLE JURISDICTION

VI. AUTHORITY

The authority is numbered

The difficulty described in the first part has a counterpart, and it is the reason I think the Brazilian set is worth building before others.

Brazilian legal authority is numbered. Cases carry a unified national docket number whose internal structure encodes court, year and origin, with check digits. Binding summaries are numbered. Repetitive themes are numbered. Statutory provisions are cited by article, paragraph and item.

Citation accuracy is therefore machine-verifiable to a degree unavailable in jurisdictions where authority is cited by party name and reporter. A fabricated docket number fails a check digit. A fabricated theme number fails a lookup against the register. There is no interpretive argument to have about it — and in a benchmark, what admits no interpretive argument is exactly what can be scored strictly.

VERIFIABILITY OF THE CITATION SURFACE

0001234-47.2019.8.26.0100

Well-formed: check digits resolve, court and origin codes exist.

0001234-99.2019.8.26.0100

Fails the check digit. Detectable without reading a word of the filing.

Tema 2.481 / STJ

Plausible in form, absent from the register. The characteristic hallucination.

This is why the argument that errors of factual detail are no better than hallucinations carries particular force here. In a Brazilian filing a transposed docket number does not weaken the piece. It attaches it to the wrong proceeding. Grading that as a criterion of medium importance would not survive contact with anyone who has filed one.

VII. THE RUBRIC

The quality standard is already codified

Legal benchmarks spend considerable effort constructing rubrics to assess reasoning quality, and the fragility of those rubrics is the most persistent methodological criticism against them: merged criteria, arguable weights, model judges whose agreement with experts goes unreported.

Brazilian law solved that problem on its own, and before AI. The Civil Procedure Code of 2015 enumerates, in six items, the circumstances in which a judicial decision is not considered reasoned. This is not doctrine or best practice: it is a defect going to the validity of the act.

Two qualifications, for precision. The list is illustrative rather than exhaustive. And the Superior Court of Justice has narrowed item VI to summaries and precedents that are formally binding, excluding merely persuasive ones — "case law" presupposes a multiplicity of decisions, and an isolated judgment is not a precedent for this purpose. Neither weakens the provision's use as an evaluation criterion. If anything the opposite: there is an enumerated, legally enforceable floor, and the case law interpreting it is itself a source of calibration.

CIVIL PROCEDURE CODE, ART. 489, § 1 — SUMMARY OF THE SIX ITEMS

Merely citing, reproducing or paraphrasing a legal provision without explaining its relation to the case.

Employing indeterminate legal concepts without explaining the concrete reason for their application.

Invoking grounds that would serve equally to justify any other decision.

Failing to address every argument capable of undermining the conclusion reached.

Invoking a precedent or binding summary without identifying its determining grounds or demonstrating that the case fits them.

Departing from a precedent or binding summary raised by a party without demonstrating distinction or overruling.

Original Portuguese text is authoritative — Law 13.105/2015, art. 489, § 1º.

Consider what items V and VI describe in the language of evaluation. Item V is the prohibition on decorative citation: invoking authority without showing why it applies. Item VI is the prohibition on selective omission: ignoring contrary authority raised by the other side. Both are precisely the failure modes legal AI systems exhibit most often, and both are verifiable by comparing the filing against the register.

Brazil does not need anyone to invent a rubric for reasoning quality. It is in the statute, enumerated, and has bound judges since 2016.

No common law jurisdiction offers an equivalent. It is a considerable methodological advantage, and it is available to whoever wants to use it.

VIII. CORPUS AND RULE

The corpus exists, and so does the regulation

Two conditions that normally have to be built from scratch are already in place.

The first is the reference register. The national judicial training school maintains, with the support of the Superior Court of Justice, a public base consolidating the binding decisions of both high courts, organised by the statutory provision each one construes and named after the article of the Civil Procedure Code that enumerates what binds. There is therefore an official register against which one can check whether a cited authority exists and what it actually decided — the input without which no citation verification is possible.

Two qualifications for anyone building on it. The register has a temporal boundary, and a task set anchored to it inherits the staleness described in the third section; dating the items is a requirement rather than a refinement. And point-by-point lookup is not the same as bulk availability: what the register guarantees is verification, not a corpus ready for download.

The second contradicts the assumption that Latin America is unregulated territory. Resolution 615 of the National Council of Justice, of 11 March 2025, has been in force since 14 July 2025 and reaches the entire judiciary with the exception of the Supreme Federal Court. It replaces and extends Resolution 332/2020, classifies uses of AI in the judiciary by risk level, requires effective and periodic human supervision so that no decision is taken exclusively by machine, imposes algorithmic governance, and created a National Artificial Intelligence Committee with representation from the judiciary, the bar, the public prosecution service, the public defender's office and civil society.

The chronology is worth recording. Brazil has had binding regulation of judicial AI in force since July 2025. The high-risk obligations of the European AI Act become applicable in December 2027.

The implication for evaluation is direct: there exists in Brazil an explicit normative set against which a system's behaviour can be measured — human supervision, traceability, handling of material under judicial secrecy. That is not a benchmark, but it is the specification from which one is written.

PART THREE: WHY NOW, AND WHAT IT WOULD TAKE

IX. THE WINDOW

Literacy, by reflex

Regulation (EU) 2026/1744 entered into force on 27 July 2026 and replaced the AI Act's article on AI literacy in its entirety. Providers and deployers no longer have to ensure a sufficient level of literacy; they must take measures to support its development. An obligation of effort, not of result.

The detail almost every analysis passed over is that the definition was left untouched. AI literacy is still defined as the skills and understanding that allow informed deployment and awareness of the risks and possible harm. The guarantee of result was removed; the standard against which the effort is measured was preserved intact.

Something counterintuitive follows. An obligation of effort is not discharged by a certificate: it is discharged by evidence. And evidence that deployment is informed does not come from training material, which ages and does not accompany the user to the moment of decision. It comes from the surface of the product — from what the system exposes about provenance, scope and uncertainty at the point where someone acts.

The vendor is not the direct addressee of that obligation when selling outside the Union, and I do not argue that it is. The effect is reflexive: whoever deploys must document their own effort, and can only document it with what the vendor supplies. In a jurisdiction where nobody has measured anything, there is nothing to document.

The article does not require a local benchmark. It makes the absence of one documentable — and since 27 July, "we don't know how the system behaves in your jurisdiction" has stopped being a commercial gap.

X. PROPOSAL

What a Brazilian task set would have to do

Nothing above argues that global benchmarks are wrong. It argues that they are not transferable, and that a jurisdiction responsible for an enormous share of the world's pending litigation — one that also holds the best technical conditions for evaluation I have seen documented anywhere — has no instrument of its own.

Tasks drawn from mass litigation as well as advisory work, weighted toward the volume where the market actually operates.

Every item dated, and scored against the law as it stood at the relevant moment — with temporal modulation treated as a first-class variable, not an edge case.

Forum recorded per item, and regionally divergent answers scored as correct within their forum rather than against a single national key.

Citation accuracy graded strictly and separately, exploiting the numbered structure of Brazilian authority. Structural validity and register existence are objective checks and should carry their own score.

Reasoning quality graded against the six items of art. 489, § 1 of the Civil Procedure Code rather than a rubric invented for the occasion — with particular attention to items V and VI, which describe the failure modes typical of current systems.

Source records that reflect the real corpus: multiple court systems, mixed digital and scanned material, inconsistent internal ordering.

Adherence to CNJ Resolution 615/2025 measured explicitly, above all on human supervision and the handling of material under judicial secrecy.

Held privately, contamination-resistant, and administered by someone with no product in the ranking.

The last point decides whether the others matter. A vendor cannot credibly grade itself in a market it is entering, and the vendors are arriving now.

XI. CODA

Before the rules are written elsewhere

Legal AI has come to be evaluated at the level of the product rather than the model, which is the correct level. What has not happened is the extension of that evaluation to jurisdictions where the underlying legal reasoning does not resemble the one the task sets were built around.

In the Brazilian case the omission is particularly hard to justify. This is not an opaque jurisdiction, or a small one, or one without rules. It is the jurisdiction that offers numbered authority, a statutory rubric for reasoning, an official reference corpus and AI regulation already in force — and that nonetheless remains without an instrument of evaluation.

Brazil will acquire a de facto standard for judging legal AI over the next two years. The question is whether it is written by people who have practised inside the system, or inherited from a task set assembled for a different one.

---

SOURCES CITED

  • Constitution of 1988, art. 93, IX — duty to give reasons for judicial decisions.
  • Law 13.105/2015 (Civil Procedure Code), art. 489, § 1, items I to VI — circumstances in which a decision is not considered reasoned.
  • CNJ Resolution 615 of 11 March 2025 — development, use and governance of AI solutions in the judiciary. In force since 14 July 2025; replaces CNJ Resolution 332/2020.
  • Law 13.709/2018 (General Data Protection Law).
  • Binding summary 203 of the Superior Court of Justice — special appeal does not lie against decisions of the second-instance panels of the small claims courts. Binding summary 640 of the Supreme Federal Court — constitutional appeal does lie against such decisions.
  • Regulation (EU) 2024/1689 (AI Act), art. 3(56) — definition of AI literacy, unamended.
  • Regulation (EU) 2026/1744 (Digital Omnibus on AI) — published in the Official Journal on 24 July 2026, in force since 27 July 2026; replaces art. 4 of the AI Act in its entirety.

About the author

Marcus Camargo

Marcus Camargo

Marcus Camargo has practised law in Brazil for thirty years and built legal data infrastructure alongside it for two decades, including the indexing of more than 285 million Brazilian court records.