Skip to content
LegalizeAI

AI Hallucination Sanctions: What Courts Are Doing

By Mark Fulton14 min read

AI Hallucination Sanctions: What Courts Are Doing

Courts are no longer treating fabricated AI citations as an embarrassing novelty. The public AI Hallucination Cases database maintained by researcher Damien Charlotin listed 1,934 decisions as of 19 August 2026, up from 16 in all of 2023, and 1,326 of them come from United States courts. Sanctions have moved from a symbolic $5,000 fine in the first well known case to five and six figure orders: the largest single US entry in the database is $110,204 in costs and penalties, and 150 entries carry a disciplinary referral. But the database does not support the conclusion vendors draw from it. A specific AI tool is named in only 219 of the 1,935 rows, and the tracker itself cautions that naming a tool does not mean that tool caused the hallucination. The verified pattern is about verification workflow, not brand choice.

We do not sell any of the tools in this directory, we take no vendor money for rankings, and featured listings are labeled and never reorder results. That matters on this topic more than any other, because sanctions stories have become the single most reliable sales asset in legal AI marketing. This page is industry analysis and general information about software, not legal advice. Professional duties differ by jurisdiction and by court, several courts now have their own standing orders on AI disclosure, and a licensed attorney in your jurisdiction is the right person to advise on any specific matter.

How big is the hallucination problem in court filings?

Big enough to have its own research infrastructure, and growing in a shape that is worth understanding precisely.

The Charlotin database is the closest thing the field has to a census. It tracks decisions in which a court or tribunal explicitly found, or implied, that a party relied on hallucinated material. It does not track every fake citation ever filed, only the ones that produced a written judicial response. Here is what the growth curve actually looks like, counted from the database's own downloadable dataset:

Quarter Decisions recorded
Q2 2023 6
Q3 2023 3
Q4 2023 7
Q1 2024 10
Q2 2024 5
Q3 2024 19
Q4 2024 25
Q1 2025 55
Q2 2025 123
Q3 2025 266
Q4 2025 401
Q1 2026 444
Q2 2026 406
Q3 2026 (through 19 August) 164

Two things stand out. The first is the scale of the jump through 2025, roughly a fourteen fold increase over the year. The second is that the curve has stopped bending upward. Q4 2025, Q1 2026 and Q2 2026 sit at 401, 444 and 406. That is a plateau, not an explosion, and it is the number to quote when someone tells you the problem is still doubling. Charlotin made the same observation to Scientific American in May 2026, describing a levelling off at roughly 350 to 400 decisions a quarter.

The composition matters more than the total. Of the 1,326 US entries, 790 involve a self represented litigant and 510 involve a lawyer. So the majority of American entries are not law firm failures at all. They are people without counsel using a free chatbot as a substitute for a lawyer, which is a real access to justice story and a completely different problem from a firm's research workflow. Anyone quoting the headline count as a measure of professional risk is inflating it by roughly half.

Within the profession, the numbers are still serious and still concentrated in a familiar failure mode. Across the whole database, 1,739 entries involve fabricated, misquoted or misrepresented case law specifically. Fabrication is the largest category at 1,609 entries, false quotations account for 522, and misrepresented authority for 806. Those categories overlap, because a single filing can do all three.

What have courts actually done about it?

They started with modest fines and a lot of patience. They have run out of the patience.

The escalation is easiest to see as a timeline. Every row below is drawn from the public database, with the court, date and outcome as the database records them.

When Case Court Tool recorded Outcome
June 2023 Mata v. Avianca, Inc. S.D.N.Y. ChatGPT $5,000 fine against two lawyers and their firm, plus letters to the client and to the judges named in the fake opinions
November 2023 Zachariah Crabill disciplinary case Colorado Supreme Court ChatGPT 90 day actual suspension, plus a stayed term and probation
January 2024 Park v. Kim 2d Cir. ChatGPT Referral to the court's grievance panel and an order to disclose the misconduct to the client
January 2025 Kohls v. Ellison D. Minn. GPT-4o Expert declaration excluded. The hallucinated citations were in an expert's filing, not a lawyer's brief
February 2025 Wadsworth v. Walmart D. Wyo. Firm internal tool $5,000 in total fines and pro hac vice admission revoked for the drafting attorney
May 2025 Lacey v. State Farm General Insurance C.D. Cal. CoCounsel, Westlaw Precision, Google Gemini $31,100 jointly and severally against two law firms, briefs struck, discovery relief denied
July 2025 Coomer v. Lindell / MyPillow, Inc. D. Colo. Several, including Copilot, Westlaw's AI, Gemini, Grok, Claude, ChatGPT and Perplexity $6,000 in monetary sanctions under Rule 11
September 2025 Noland v. Land Cal. Ct. App. Unidentified $10,000, State Bar notified, opinion ordered served on the client. The court counted 21 fabrications among 23 case quotations in the opening brief
March 2026 Whiting v. City of Athens, Tenn. 6th Cir. Implied $30,000 in sanctions and adverse costs, plus potential disciplinary proceedings
March 2026 Couvrette v. Wisnovsky D. Or. Unidentified Briefs struck, $15,500 sanction, $94,700 adverse costs order, claims dismissed with prejudice
August 2026 Kleyman Law Group, P.C. v. Kaloidis N.Y. Sup. Ct. Claude and other generative AI $46,511

Read down that column and the trend is clear enough to plan around. In 2023 and 2024 the database records six decisions carrying a monetary penalty and five carrying a professional sanction. In 2025 alone it records 182 monetary penalties and 83 professional sanctions, and 2026 through mid August already shows 163 and 62. Across the whole database, 351 entries carry a monetary penalty. Of those, 192 are denominated in US dollars, and they total just over $1.28 million.

The severity ceiling has risen too. Mata v. Avianca cost $5,000 in 2023. Couvrette v. Wisnovsky cost $110,204 across sanctions and costs in 2026, and it is worse than that number suggests, because the claims were dismissed with prejudice after the offending briefs were struck. The client lost the case partly because the brief was unusable.

The legal hook is not new law. It is Federal Rule of Civil Procedure 11, which has said the same thing since long before generative AI existed: by signing and filing a paper, an attorney certifies that after a reasonable inquiry, the legal contentions are warranted by existing law. A citation that does not exist cannot be warranted by existing law, and no reasonable inquiry was made if nobody opened the case. Courts have not needed a new rule. They have simply applied the old one, and Rule 11(c) already gave them monetary sanctions, fee shifting and nonmonetary directives to work with.

Why do general chatbots keep causing this?

Because a general model is optimized to produce a plausible next token, and a citation is one of the most predictable text patterns in the language. Case names follow a template. Reporter volumes and page numbers are integers in a familiar range. Pin cites look like other pin cites. A model that has read a great deal of case law learns exactly what a citation to a case that would support this argument should look like, and it can produce one fluently whether or not the case is real.

That is why the fabrications are so convincing. In Mata, the fake decisions came complete with fabricated internal quotations. In Noland, the court found 21 fabricated quotations in a single opening brief. These are not garbled outputs that a careful reader would spot from across the room. They read like law.

The frequency is measurable, and the strongest independent measurement is still Stanford's. In a 2024 benchmark from Stanford HAI and RegLab, researchers ran a pre-registered set of more than 200 open ended legal queries through purpose built legal research products. Lexis+ AI and Ask Practical Law AI produced incorrect information more than 17% of the time. Westlaw's AI Assisted Research hallucinated more than 34% of the time. The same team's earlier work on general purpose chatbots found hallucination rates between 58% and 82% on legal queries.

Set those side by side and you get the honest version of the story. Retrieval grounding cuts the error rate substantially. It does not take it to zero, and a one in six error rate on a research question is not a rate you can file against without checking.

There is a second cause that has nothing to do with model architecture. Reporting in Scientific American in May 2026 described researchers characterising the behaviour as cognitive surrender: users under time pressure hand the thinking over completely, and people who feel most positive about AI verify least. The pattern in the case records supports this. In Wadsworth, eight of nine citations in the filed motions were nonexistent or pointed to different cases, and nobody opened them. In Lacey, a hallucinated outline was passed between colleagues at two firms and incorporated into a filed brief without anyone checking the underlying authorities. The tool produced the error. The workflow shipped it.

Which tool categories are built to prevent it?

Here is where we part company with most writing on this subject, because the honest answer is that the case database cannot rank tools, and anyone using it to do so is selling something.

Of 1,935 entries, 1,296 record the tool as merely implied and 402 as unidentified. A specific product is named in 219. Of those, ChatGPT accounts for 106, with small numbers for Copilot, Claude, Perplexity, Grok, Gemini and a handful of legal specific products. The database carries an explicit warning on that column: naming a tool does not necessarily mean that tool was responsible for the hallucinations in question. Courts record what the sanctioned party admits to, and general chatbots are both the most used and the most readily confessed. You cannot derive a safety ranking from a confession rate.

What you can do is reason about product architecture, which is a category question rather than a brand question. Four categories reduce citation risk in genuinely different ways:

Category What it does about fabrication What it does not solve
Grounded legal research platforms Answers are generated over a licensed or proprietary corpus of primary law, with citations resolved against real documents and, in some products, a citator signal for treatment Retrieval can still surface the wrong case, and the summary of a real case can still misstate its holding. Stanford measured this directly
Citation verification and brief checking tools Run over a finished draft and check every citation against the record or against primary law before it is filed. This is the last line of defense and the only one that catches an error introduced anywhere upstream Only as good as its coverage of the sources you cite, and it does not stop bad reasoning, only bad references
Drafting assistants with source linking Every generated assertion carries a link back to the passage it came from, so verification is a click instead of a research project The link tells you the passage exists. Whether it supports the proposition is still your judgment
General purpose chatbots Nothing structural. Some now search the web, which changes the failure mode from invention to misreading rather than removing it Everything above. These are drafting and thinking tools, not sources of authority

The verification layer is the one buyers underweight, and it is the one that maps most directly to the failure pattern in the case records. A tool like Clearbrief, which sits inside Word and checks citations against the record before a brief goes out, is addressing the exact step that was skipped in Wadsworth and Lacey. You can see the full set of grounded research platforms on our legal research category page, where every listing states what it is built on.

Note also who is not in the database in the way you would expect. Of the 1,326 US entries, most involve people with no lawyer at all. No amount of firm procurement fixes that half of the problem, which is one reason the question of whether AI can replace a lawyer has an increasingly evidence backed answer.

What does this mean for how you evaluate tools?

Four things, and none of them is "buy the tool from the vendor that published the scariest sanctions blog post."

Treat grounding as a claim to test, not a feature to check off. Every research vendor now says its citations are grounded in primary law. Stanford's benchmark found meaningful differences between products that all made that claim. Ask for the error rate on your own practice area, run a pilot on questions where you already know the answer, and count the failures yourself. Our buyer's checklist for evaluating legal AI has the twenty questions to put in writing.

Buy verification separately from generation. These are different jobs, and the market prices them differently. A grounded research platform reduces how often an error appears. A citation checker catches the errors that appear anyway, including the ones a paralegal, a co-counsel, an expert or a contract drafter introduced. The case records include hallucinations that entered through an expert declaration and through a freelance drafter, neither of whom your research subscription covers.

Weight the workflow over the model. In almost every sanctioned filing, the identifiable failure is that no human opened the cited authority before it was filed. That is a process control, and it costs nothing. A rule that every citation in a filed document has been opened and read by a named person will do more for your risk profile than any procurement decision.

Discount the fear marketing, and check the date on any count you are quoted. Numbers in this field go stale in weeks. A page citing a figure from six months ago is describing a database that has since grown by several hundred entries, and the growth rate itself has changed shape. If a vendor's case count does not carry an as of date, it is decoration.

The trend line here is not that AI is unusable in legal work. It is that courts have settled on a standard, they applied an existing rule to reach it, and the standard is the same one that applied to a junior associate with a photocopier: you certify what you file. The tools that help are the ones that make certification fast rather than the ones that promise you will not need to.

Frequently asked questions

Have lawyers been sanctioned for using ChatGPT?

Yes, repeatedly, though the precise framing matters. Courts sanction the failure to verify, not the use of the tool. In Mata v. Avianca in June 2023, two lawyers and their firm were fined $5,000 after filing a brief containing nonexistent decisions produced by ChatGPT. In November 2023 the Colorado Supreme Court imposed a 90 day actual suspension on attorney Zachariah Crabill for filing ChatGPT generated citations he had not read. In the public database, ChatGPT is the named tool in 106 entries, more than any other product, although the tool goes unrecorded in the large majority of cases. Duties and available sanctions vary by jurisdiction and by court, and several courts now have standing orders on AI use.

How many AI hallucination court cases are there?

As of 19 August 2026, the public AI Hallucination Cases database listed 1,934 decisions worldwide, of which 1,326 are from US courts. The next largest jurisdictions are Canada with 211, Australia with 98, the UK with 62 and Israel with 55. The count grew from 16 decisions in all of 2023 to 845 in 2025, and 2026 has recorded over 1,000 through mid August. That number changes weekly, so check the database rather than any article, including this one, for the current figure.

Do legal research AIs still hallucinate?

Yes, at lower rates than general chatbots. The most rigorous independent measurement remains Stanford's 2024 benchmark, which found Lexis+ AI and Ask Practical Law AI producing incorrect information more than 17% of the time and Westlaw's AI Assisted Research hallucinating more than 34% of the time on over 200 open ended legal queries, against 58% to 82% for general purpose chatbots on legal questions. Vendors have shipped a great deal since that study, and none of them publishes an independently audited current figure, which is itself worth noting. Lacey v. State Farm is the case to remember here: the hallucinated material came out of purpose built legal AI products, not a consumer chatbot.

What is citation grounding in legal AI?

Grounding means the system does not generate a citation from its own parameters. It retrieves actual documents from a corpus of primary law, then generates its answer over those retrieved documents and cites the ones it used, so every citation points at a record that exists. Strong implementations go further: they link each assertion to the specific passage it came from, resolve the citation against a citator so you can see whether the case is still good law, and refuse to answer when retrieval finds nothing. Grounding reduces invention. It does not guarantee the retrieved case supports the proposition, which is why the verification step stays yours.

Every tool in our legal research directory lists what corpus it is built on and how it handles citations, with no paid placement affecting the order. Start there if you are picking a research platform, and pair whatever you choose with a verification step before anything gets filed.