Tasks, claims, datasets, splits, leakage, contamination, baselines, metrics, calibration, robustness, citation support, factuality, abstention, subgroup slices, human review, misuse, monitoring, and honest reporting.

Structured Visual

Jurisdiction: US; as of 2026-08-28; not legal advice; Render structure, refuse interpretation, cite, abstain, and hand off.

RENDER STRUCTURE · REFUSE INTERPRETATION · CITE · ABSTAIN · HAND-OFF: render structure, refuse interpretation, cite provenance, abstain when unsupported, and hand off to human review.

Evaluating Legal AI: selected questionsSelected questionsClaim and taskDataset governanceLeakage
highlighted = computed this step

Scope and honesty note

Jurisdiction: United States computational-law classroom model; source snapshot 2026-08-28; curriculum as of 2026-08-29. Synthetic inputs, code, labels, measurements and outputs are teaching artifacts, not law, legal advice, authority, eligibility, benefits, tax, court, filing, research, ranking, or outcome determinations. Code encodes selected interpretations and can be incomplete, wrong, outdated, biased, overprecise, underinclusive, or non-isomorphic. The system must cite source and version, expose assumptions and gaps, abstain when unsupported, and hand legal judgment to accountable humans.

computational-law snapshot 2026−08−28\text{computational-law snapshot }2026-08-28

See the essential structure first

Start with this deliberately incomplete structure, then use the pinned authorities, worked application, exceptions, and handoff below. This deliberately incomplete preview has 4 nodes; exceptions and legal consequences remain in the sourced prose below.

glance nodes=4\text{glance nodes}=4

Jurisdiction: US; as of 2026-08-28; not legal advice; Render structure, refuse interpretation, cite, abstain, and hand off.

RENDER STRUCTURE · REFUSE INTERPRETATION · CITE · ABSTAIN · HAND-OFF: render structure, refuse interpretation, cite provenance, abstain when unsupported, and hand off to human review.

Evaluating Legal AI: selected questionsSelected questionsClaim and taskDataset governanceLeakage

Begin with computational-law doctrine

Legal-AI evaluation begins with a precise task and claim, not a leaderboard. Datasets require provenance, licenses, annotation, deduplication, temporal splits and leakage audits. Metrics must match use: retrieval ranking, extraction spans, calibration, claim-to-source support, citation correctness, factuality and abstention are distinct. Aggregate scores can hide jurisdictional, temporal, OCR, ambiguity and subgroup failures. Human evaluation needs rubrics and disagreement records. No numeric benchmark result appears without a measured, reproducible experiment. Deployment requires monitoring and reevaluation, not a one-time pass.

source, semantics, trace, uncertainty, human judgment\text{source, semantics, trace, uncertainty, human judgment}

Bounded evaluation sample

The three-record valid sample illustrates dataset cards, source hashes, temporal splits and quarantine without becoming a benchmark score. Pinned source or measurement: “{"case_count": 1534, "cases": [{"case_path": "/us/410/0113-01", "date": "1973-01-22", "importance": 0.003491912312537867, "name": "Roe v. Wade", "sections": ["28 U.S.C. § 1253"], "url": "catalog/cases/us/volume_410/0113_01/index.html"}, {"case_path": "/us/501/0722-01", "date": "1991-06-24", "importance": 0.003330523929892698, "name": "Coleman v. Thompson", "sections": ["28 U.S.C. § 1257"], "url": "catalog/cases/us/volume_501/0722_01/index.html"}, {"case_path": "/us/501/0808-01", "date": "1991-06-27", "importance": 0.002951839720053311, "name": "Payne v. Tennessee", "sections": ["28 U.S.C. § 2281"], "url": "catalog/cases/us/volume_501/0808_01/index.html"}], "key": "usc_title_28", "label": "Judiciary and Judicial Procedure", "title": 28}” Coordinate: cases_by_law validity-filtered Title 28 sample as of 2026-08-28; https://www.neochart.com/cases-by-law/; data via neochart.com, snapshot 2026-08.

pinned coordinate: casesbylawvalidity−filteredTitle28sampleasof2026−08−28\text{pinned coordinate: }cases_by_law validity-filtered Title 28 sample as of 2026-08-28

Evidence-linked relation sample

The bounded edge set supports citation-extraction and graph-query evaluation with exact expected relations. Pinned source or measurement: “[{"from": "/f-supp-2d/692/0170-01", "fromDate": "2010-03-09", "fromName": "Watkins v. Omni Life Science, Inc.", "to": "/us/559/0077-01", "type": "cites"}, {"from": "/f-supp-2d/712/0924-01", "fromDate": "2010-05-17", "fromName": "In re Arrowhead Capital Management LLC Class Litigation", "to": "/us/559/0077-01", "type": "cites"}, {"from": "/f-supp-2d/718/0805-01", "fromDate": "2010-03-10", "fromName": "Astra Oil Trading NV v. Petrobras America Inc.", "to": "/us/559/0077-01", "type": "cites"}]” Coordinate: case-citations-hertz-inbound-slice as of 2026-08-28; https://www.neochart.com/catalog/cases/us/volume_559/0077_01/index.html; data via neochart.com, snapshot 2026-08.

pinned coordinate: case−citations−hertz−inbound−sliceasof2026−08−28\text{pinned coordinate: }case-citations-hertz-inbound-slice as of 2026-08-28

High-stakes claim source

The statute provides exact propositions for citation-support evaluation and exposes harm from unsupported legal summaries. Pinned source or measurement: “§706. Scope of review To the extent necessary to decision and when presented, the reviewing court shall decide all relevant questions of law, interpret constitutional and statutory provisions, and determine the meaning or applicability of the terms of an agency action. The reviewing court shall- (1) compel agency action unlawfully withheld or unreasonably delayed; and (2) hold unlawful and set aside agency action, findings, and conclusions found to be- (A) arbitrary, capricious, an abuse of discretion, or otherwise not in accordance with law; (B) contrary to constitutional right, power, privilege, or immunity; (C) in excess of statutory jurisdiction, authority, or limitations, or short of statutory right; (D) without observance of procedure required by law; (E) unsupported by substantial evidence in a case subject to sections 556 and 557 of this title or otherwise reviewed on the record of an agency hearing provided by statute; or (F) unwarranted by the facts to the extent that the facts are subject to trial de novo by the reviewing court. In making the foregoing determinations, the court shall review the whole record or those parts of it cited by a party, and due account shall be taken of the rule of prejudicial error. ( Pub. L. 89–554, Sept. 6, 1966, 80 Stat. 393 .)” Coordinate: 5 U.S.C. § 706; https://www.neochart.com/catalog/federal/title_5/section_706/title5_sec706_9a580d7b5bc7/706_scope_of_review_to_the_extent_necessary_to_decision_and_0001/index.html; data via neochart.com, snapshot 2026-08.

pinned coordinate: 5U.S.C.§706\text{pinned coordinate: }5 U.S.C. § 706

Pin the synthetic computational record

A synthetic evaluation plan uses the pinned statutes, valid case sample and Hertz edges to populate task cards, dataset manifests, annotation and split records, leakage candidates, baselines, metric definitions with intentionally blank result values, robustness and slice cases, citation-support rubrics, abstention cases, human review, errors, limitations, deployment gates and monitoring.

stated inputs and operations, not legal conclusions\text{stated inputs and operations, not legal conclusions}

Work the audited application

The citation parser is evaluated against the three exact Hertz edges, while statute QA requires each atomic claim to match section Seven-Zero-Six text. A random split is rejected because duplicate and later opinions leak citation language; a temporal, cluster-aware split is chosen. Unsupported questions test abstention. Since no experiment runs in the lesson, numeric result and improvement fields remain empty rather than invented.

execute, explain, test, abstain, hand off\text{execute, explain, test, abstain, hand off}

Read the populated computational artifact

The evaluation record contains system, claim, user, context, task, input, output, prohibited use, dataset, source, license, snapshot, annotation, guideline, adjudication, split, temporal cutoff, duplicate cluster, leakage type, contamination, baseline, metric definition, result field, interval field, calibration, claim support, citation correctness, abstention, slice, stress case, subgroup, human rubric, reviewer, disagreement, error, limitation, misuse, drift, incident, rollback, and reevaluation. The artifact contains 17 populated rows.

rows=17\text{rows}=17

Jurisdiction: US; as of 2026-08-28; not legal advice; Render structure, refuse interpretation, cite, abstain, and hand off.

RENDER STRUCTURE · REFUSE INTERPRETATION · CITE · ABSTAIN · HAND-OFF: render structure, refuse interpretation, cite provenance, abstain when unsupported, and hand off to human review.

Evaluating Legal AI: Pinned sources and measurementsPinned sources and measurementsVerbatim text or bounded…cases_by_law validity-filtered Title 28 sample as of 2026-08-28: Bounded evaluation sampleThe three-record valid sample…case-citations-hertz-inbound-slice as of 2026-08-28: Evidence-linked relation sampleThe bounded edge set…5 U.S.C. § 706: High-stakes claim sourceThe statute provides exact…
Evaluating Legal AI: Synthetic computational recordSynthetic computational recordClassroom inputs and intermediate…System claimsResearch retrieval, citation extraction,…DatasetDated source snapshot, annotation…ReportBaselines, metric definitions, unreported…
Evaluating Legal AI: Audit trace part 1Audit traceSemantics, provenance, execution, evidence,…Claim and taskIntended user, decision context,…Dataset governanceSource, license, provenance, snapshot,…LeakageSame document, parallel opinion,…
Evaluating Legal AI: Audit trace part 2Audit traceSemantics, provenance, execution, evidence,…MetricsRetrieval recall and ranking,…Slices and stressJurisdiction, court, date, document…Human evaluationRubric, qualification, blinding, disagreement,…
Evaluating Legal AI: Audit trace part 3Audit traceSemantics, provenance, execution, evidence,…Reporting and monitoringNo invented numbers, baseline,…

Read the complete record

The complete record keeps sources, stated facts, and questions for review separate. Pinned sources and measurements: Verbatim text or bounded snapshot data. cases_by_law validity-filtered Title 28 sample as of 2026-08-28: Bounded evaluation sample: The three-record valid sample illustrates dataset cards, source hashes, temporal splits and quarantine without becoming a benchmark score.. case-citations-hertz-inbound-slice as of 2026-08-28: Evidence-linked relation sample: The bounded edge set supports citation-extraction and graph-query evaluation with exact expected relations.. 5 U.S.C. § 706: High-stakes claim source: The statute provides exact propositions for citation-support evaluation and exposes harm from unsupported legal summaries.. Synthetic computational record: Classroom inputs and intermediate states. System claims: Research retrieval, citation extraction, statute QA, rule execution and abstention are evaluated as separate tasks. Dataset: Dated source snapshot, annotation guidelines, unit of analysis, train or validation or test split, duplicates, temporal cutoff and access controls. Report: Baselines, metric definitions, unreported numeric fields, slices, confidence intervals, errors, red-team cases, human review, limitations and deployment monitoring. Audit trace: Semantics, provenance, execution, evidence, limits and handoff. Claim and task: Intended user, decision context, input, output, allowed assistance, prohibited use, unit, ground truth, consequence and acceptance criterion. Dataset governance: Source, license, provenance, snapshot, annotation, adjudication, split, duplication, contamination, benchmark exposure and change history. Leakage: Same document, parallel opinion, quote overlap, citation neighborhood, temporal future, template, entity, label leakage and memorized benchmark. Metrics: Retrieval recall and ranking, span precision or recall, classification calibration, claim support, citation correctness, factuality, abstention quality, latency and cost. Slices and stress: Jurisdiction, court, date, document type, OCR quality, rare label, ambiguity, conflicting authority, missing source, adversarial query and subgroup impact. Human evaluation: Rubric, qualification, blinding, disagreement, adjudication, fatigue, conflicts, qualitative errors and inter-reviewer measurement. Reporting and monitoring: No invented numbers, baseline, interval, significance limits, error examples, limitations, misuse, drift, incidents, feedback, rollback and reevaluation.

sources, stated facts, and open questions\text{sources, stated facts, and open questions}

Narrow summary

Evaluate precise claims on provenance-rich leakage-audited data, match metrics to tasks and harms, report slices and errors, invent no scores, and monitor after deployment.

trace, test, abstain, hand off\text{trace, test, abstain, hand off}