Testing the answer against law, evidence and source material
The central question is whether an answer is supported by the law and the source material, not simply whether it is fluently expressed.
Model-response evaluation
Assessment of competing legal answers for accuracy, quality of reasoning, support from authority, completeness and professional usefulness.
Hallucination and authority checking
Checking cited cases, legislation and legal propositions against the source material, and identifying claims that go beyond what the sources establish.
Rubrics and benchmarks
Designing and applying realistic evaluation criteria, adversarial legal scenarios and tasks that reveal weaknesses which fluent drafting may conceal.
Long-context evaluation
Testing whether outputs accurately reflect lengthy records, preserve important distinctions and avoid unsupported assumptions.
Testing an answer against the record
Legal work often turns on small but decisive distinctions between what a document states, what can properly be inferred from it and what remains unproved. Evaluation includes tracing important conclusions back to the source material, identifying omitted qualifications and testing the reasoning against the underlying record.
Legal research and drafting
Assessment of AI-assisted legal research, fact-checking, source grounding and drafting, together with concise expert feedback and red-team testing of difficult legal questions.
Professional usefulness
Review of whether outputs are appropriately qualified, sufficiently complete for the task and suitable for use within a professional legal workflow.
Legal technology & domain consultancy
Where a project extends beyond model evaluation to the design of legal tasks, workflows or products, the related consultancy service provides broader legal-domain input.