Classification, NER, and Extraction
Classification, NER, and Extraction over Opinions
Label design, annotation, spans, entities, relations, events, holdings, procedural posture, citations, sections, weak supervision, models, evaluation, disagreement, OCR noise, domain shift, and honest errors.
Structured Visual
Jurisdiction: US; as of 2026-08-28; not legal advice; Render structure, refuse interpretation, cite, abstain, and hand off.
RENDER STRUCTURE · REFUSE INTERPRETATION · CITE · ABSTAIN · HAND-OFF: render structure, refuse interpretation, cite provenance, abstain when unsupported, and hand off to human review.
Scope and honesty note
Jurisdiction: United States computational-law classroom model; source snapshot 2026-08-28; curriculum as of 2026-08-29. Synthetic inputs, code, labels, measurements and outputs are teaching artifacts, not law, legal advice, authority, eligibility, benefits, tax, court, filing, research, ranking, or outcome determinations. Code encodes selected interpretations and can be incomplete, wrong, outdated, biased, overprecise, underinclusive, or non-isomorphic. The system must cite source and version, expose assumptions and gaps, abstain when unsupported, and hand legal judgment to accountable humans.
See the essential structure first
Start with this deliberately incomplete structure, then use the pinned authorities, worked application, exceptions, and handoff below. This deliberately incomplete preview has 4 nodes; exceptions and legal consequences remain in the sourced prose below.
Jurisdiction: US; as of 2026-08-28; not legal advice; Render structure, refuse interpretation, cite, abstain, and hand off.
RENDER STRUCTURE · REFUSE INTERPRETATION · CITE · ABSTAIN · HAND-OFF: render structure, refuse interpretation, cite provenance, abstain when unsupported, and hand off to human review.
Begin with computational-law doctrine
Legal classification and extraction depend first on a defensible task and label ontology. Annotation guidelines, evidence spans, unknown and ambiguity states, disagreement and adjudication are part of the dataset, not noise to hide. OCR, quotations, footnotes, nested entities, aliases, procedural posture and domain shift make opinions messy. Evaluation must match downstream use and report slices and uncertainty. This chapter invents no benchmark numbers. Extracted labels are probabilistic data products and cannot establish holdings, facts, credibility or case outcomes without source review.
Real opinion sample
The three valid records supply names, dates, sections, paths and measured fields for bounded annotation. Pinned source or measurement: “{"case_count": 1534, "cases": [{"case_path": "/us/410/0113-01", "date": "1973-01-22", "importance": 0.003491912312537867, "name": "Roe v. Wade", "sections": ["28 U.S.C. § 1253"], "url": "catalog/cases/us/volume_410/0113_01/index.html"}, {"case_path": "/us/501/0722-01", "date": "1991-06-24", "importance": 0.003330523929892698, "name": "Coleman v. Thompson", "sections": ["28 U.S.C. § 1257"], "url": "catalog/cases/us/volume_501/0722_01/index.html"}, {"case_path": "/us/501/0808-01", "date": "1991-06-27", "importance": 0.002951839720053311, "name": "Payne v. Tennessee", "sections": ["28 U.S.C. § 2281"], "url": "catalog/cases/us/volume_501/0808_01/index.html"}], "key": "usc_title_28", "label": "Judiciary and Judicial Procedure", "title": 28}” Coordinate: cases_by_law validity-filtered Title 28 sample as of 2026-08-28; https://www.neochart.com/cases-by-law/; data via neochart.com, snapshot 2026-08.
Relation extraction sample
The three evidence-linked edges provide a citation-relation target distinct from doctrinal relation labels. Pinned source or measurement: “[{"from": "/f-supp-2d/692/0170-01", "fromDate": "2010-03-09", "fromName": "Watkins v. Omni Life Science, Inc.", "to": "/us/559/0077-01", "type": "cites"}, {"from": "/f-supp-2d/712/0924-01", "fromDate": "2010-05-17", "fromName": "In re Arrowhead Capital Management LLC Class Litigation", "to": "/us/559/0077-01", "type": "cites"}, {"from": "/f-supp-2d/718/0805-01", "fromDate": "2010-03-10", "fromName": "Astra Oil Trading NV v. Petrobras America Inc.", "to": "/us/559/0077-01", "type": "cites"}]” Coordinate: case-citations-hertz-inbound-slice as of 2026-08-28; https://www.neochart.com/catalog/cases/us/volume_559/0077_01/index.html; data via neochart.com, snapshot 2026-08.
Pin the synthetic computational record
A synthetic annotation project uses the valid cases-by-law sample and Hertz edges to populate document versions, raw and normalized offsets, entity, citation and procedural-event spans, label guidelines, annotators, disagreements, adjudication, weak labels, split provenance, model predictions, confidence, error categories and reviewer notes.
Work the audited application
The annotators distinguish a quoted court from the deciding court and a cited case from a party. A citation relation uses the pinned edge evidence. Disagreement over a holding span remains recorded until adjudication rather than majority-voted away. OCR-split names and nested statute citations receive partial-span error labels. With no actual trained run, metric fields remain unreported.
Read the populated computational artifact
The extraction record contains corpus snapshot, document, page, raw text, normalized text, offset map, task, label, definition, positive example, negative example, hierarchy, entity, span, relation, event, citation, section, posture, disposition field, annotator, guideline version, disagreement, adjudication, rationale, weak label, split, model, prediction, confidence, exact match, partial match, metric field, slice, OCR error, quotation, nested entity, alias, leakage, prohibited inference, and reviewer. The artifact contains 16 populated rows.
Jurisdiction: US; as of 2026-08-28; not legal advice; Render structure, refuse interpretation, cite, abstain, and hand off.
RENDER STRUCTURE · REFUSE INTERPRETATION · CITE · ABSTAIN · HAND-OFF: render structure, refuse interpretation, cite provenance, abstain when unsupported, and hand off to human review.
Read the complete record
The complete record keeps sources, stated facts, and questions for review separate. Pinned sources and measurements: Verbatim text or bounded snapshot data. cases_by_law validity-filtered Title 28 sample as of 2026-08-28: Real opinion sample: The three valid records supply names, dates, sections, paths and measured fields for bounded annotation.. case-citations-hertz-inbound-slice as of 2026-08-28: Relation extraction sample: The three evidence-linked edges provide a citation-relation target distinct from doctrinal relation labels.. Synthetic computational record: Classroom inputs and intermediate states. Annotation set: Three opinions with raw text, pages, OCR artifacts, citations, party and court metadata and procedural passages. Schema: Case name, party, court, judge, date, citation, statute, claim, motion, disposition, quoted rule, fact event and uncertainty spans. Experiment: Guideline version, annotators, disagreements, adjudication, train or validation or test split, model version, predictions, errors and no published score. Audit trace: Semantics, provenance, execution, evidence, limits and handoff. Task definition: Document or sentence classification, token or span NER, relation, event, citation, section segmentation, summarization input and downstream use. Labels: Operational definition, positive and negative examples, hierarchy, multilabel, unknown, ambiguous, not applicable and prohibited inference. Annotation: Raw span offsets, normalized offsets, page, evidence, annotator, guideline, timestamp, disagreement, adjudication and rationale. Models: Rules, dictionaries, CRF, transformer, prompt, weak supervision, ensemble and calibration; architecture does not cure label ambiguity. Evaluation: Exact and partial span, precision, recall, F score, confusion matrix, calibration, slice, bootstrap or interval, adjudicator agreement and error taxonomy. Messiness: OCR, hyphenation, footnotes, citations, quotations, nested entities, party aliases, procedural complexity, domain and temporal shift and leakage. Safety: No outcome, credibility, guilt, dangerousness, legal merit, protected trait or sensitive inference from opinion text.
Narrow summary
Design labels before models, preserve spans and disagreement, evaluate on honest slices without invented numbers, and verify every extracted legal claim against the opinion.