Legal AI still hallucinates in 1 in 5 to 3 in 10 searches: how to build a verification layer
Benchmarks from 2025-2026 show legal AI tools still err in 15-35% of searches. Here is how to set up a verification layer over Lexis and Westlaw.
You must establish a separate verification workflow that checks every AI-generated legal finding against primary sources before relying on it in any matter. The tools themselves remain demonstrably unreliable; the layer you build around them is what makes them safe to use.
An analysis of 10 September 2026 of how legal AI tools hallucinate in realistic searches argues that purpose-built legal AI systems still err in 15 to 35 percent of searches, even on structured statutory tasks. Recent independent benchmarking measured hallucination rates across leading legal research platforms and found that accuracy on complex tasks remains stuck at around 60 to 65 percent. In our assessment, that residual error is too large to treat an interface output as a finished authority, particularly in high-stakes work where an appeal, supervisory advice or cross-border opinion depends on the answer being correct.
Why the interface alone is not enough
Legal-specific AI tools do perform better than general models, but that advantage creates a particular risk: if a tool is demonstrably better than a junior member of staff, the temptation arises to check less. A convincingly worded answer may rest on out-of-date law. The tool cannot reliably signal when it is drawing on superseded legislation or when a citation does not actually exist. Controlled benchmarks show roughly one in five to three in ten searches produce hallucinations or incorrect answers. In daily practice, where questions vary more widely than in test conditions, that error margin may shift in some contexts.
What does a verification layer actually do?
The layer is not an additional tool but a governed workflow that makes visible how an AI finding becomes a checked authority. It separates finding from validation. You route the AI output through a documented checking process: verify each citation against primary sources, confirm the current status of any legislation cited, and record human approval before the answer leaves your control. For sensitive files, the workflow can anonymise content before sending it to any external model, and halt the process if a privacy check fails.
Which concrete controls do you need to demonstrate?
- Document the source and method — record which AI tool generated the finding, which primary sources were checked against it, and the date of that check.
- Verify citations independently — confirm that each case, statute or regulation cited actually exists and is quoted accurately.
- Check currency — establish that any legislation or case law cited remains current law and has not been superseded or overruled.
- Record human approval — document which person reviewed the finding, what they checked, and whether they approved it for use.
- Retain the audit trail — keep the original AI output, the checking steps, and the approval decision together so the workflow is traceable.
What can tooling carry and what remains your own judgement?
Automated systems can route tasks through selected models, flag inconsistencies between outputs, and make checking steps visible for inspection. They cannot make the final judgement on whether an answer is fit for purpose in your specific matter. The professional responsibility for that judgement remains with you. Tooling makes the checking visible and traceable; it does not remove the need for it. The benchmarks do not argue against using legal AI tools; they argue for discipline around how you use them. A verification layer is that discipline made operational.
Sources: This article draws on reporting and guidance from Stanford RegLab, Deliberative Democracy & Human-Centered AI, LegalAIInsights, AI Vortex and LawNext.
Written by
Marit Halversen
Covers AI governance and regulatory design, with a focus on how compliance obligations land on architecture rather than on paperwork.