The Newest AI Models Still Hallucinate Case Law, Just More Convincingly

MIT researchers found that AI models are 34% more likely to use confident, certain language, 'definitely,' 'clearly,' 'well-established', precisely when they're hallucinating

Share
The Newest AI Models Still Hallucinate Case Law, Just More Convincingly

Everyone assumes the hallucination problem is basically solved by now. Newer models, better reasoning, fewer mistakes... right?

A bankruptcy court in the Southern District of Texas would disagree.

On July 14, 2026, a firm relying on Westlaw Precision, not a free chatbot, a premium legal research tool built specifically to be more reliable, filed briefs citing three cases that don't exist and misrepresenting two more. The court didn't just strike the filings. It ordered $29,877 in sanctions, held the firm in civil contempt, and mandated continuing legal education on generative AI.

This wasn't an isolated slip. Ten days later, a lawyer in Washington state filed a brief built with both ChatGPT and Claude, running two models is often pitched as a safeguard, and still landed a $3,000 sanction for fabricated case law, invented doctrine, and misquoted authority.

Here's what should worry every practicing attorney: the fabrications in 2026 don't read like fabrications. MIT researchers found that AI models are 34% more likely to use confident, certain language, "definitely," "clearly," "well-established", precisely when they're hallucinating. OpenAI's own reporting shows its newer reasoning models hallucinate on 33–51% of open-domain factual questions, often more than the models they replaced.

The Damien Charlotin hallucination case database has now logged 1,809 cases. New entries are added almost daily. The tools didn't get more careful. They got better at sounding right.

Model upgrades are not a verification strategy. Reading every cited case remains the only one.

What's your firm's actual practice for verifying AI-assisted citations before filing?