Brand hallucination monitoring means checking on a schedule what AI models state as fact about your company, and correcting the false claims before a buyer repeats them in a meeting. This article is about answers that are wrong, like a price you do not charge, because those you can fix on a page this week. Answers that are true but unflattering are a different problem, and I will cover them in another post.
A hallucination arrives in the same fluent sentence as everything else the model says. When Claude described a pricing tier for cloudweld.ai that we have never sold, it used the same tone it used for the facts it got right, and a buyer reading that answer had no way to tell the two apart.
How often it happens
Two research groups have put numbers on how often models invent claims. I believe those numbers deserve more of your attention than any single screenshot, because the polish of a generated answer hides how often it is wrong.
Vectara's hallucination leaderboard hands a model a document and asks it to summarize only what is on the page, and a task does not get easier than that for a model. Even with the source text in front of them, the flagship OpenAI and Anthropic models invent facts in roughly 5 to 12 percent of summaries on the May 11, 2026 update of the board (HHEM-2.3), with GPT-4.1 at 5.6 percent, GPT-5.5 at 9.3, Claude Sonnet 4 at 10.3 and Claude Opus 4.5 at 10.9. The full board is wider than that band, since the best small models sit under 2 percent and some reasoning models pass 20. The leaderboard is a grounded summarization test that tells the model to use only the document and none of its internal knowledge, so I read it as a floor for questions about your brand and I would never quote it as a brand hallucination rate.
The reasoning rows on the same board are why I call it a floor. OpenAI's o3-Pro scores 23.3 percent and o4-mini-high 18.6 against GPT-4.1 at 5.6, and Vectara's January 2025 test found DeepSeek-R1 at 14.3 percent against DeepSeek-V3 at 3.9, with Vectara's own caveat that it was too early to generalize from one pair of models. A longer chain of thought gives the model more chances to state a guess as a fact, and that matches what those rows show.
The Tow Center study is closer to what buyers do, because it measured models answering real questions with retrieval switched on. In March 2025 the Tow Center for Digital Journalism at Columbia ran 1,600 queries across eight AI search engines, asking each to identify the source of a direct quote, and the engines failed more than 60 percent of the time. Perplexity was wrong 37 percent of the time and Grok-3 was wrong 94 percent of the time. ChatGPT Search got 134 of its 200 answers wrong and signaled uncertainty only 15 times across all 200 responses.
I see no reason to expect a system that misattributes a checkable newspaper quote six times out of ten to be more careful with your pricing page.
What a brand hallucination has cost
There are now enough public incidents to put a date and a price on it. My issue with the hypothetical examples in the vendor decks I have seen is that they let you treat the risk as something for next year, while the two cases below come with a tribunal decision and a public apology attached, and the two from this year are still in the news.
In February 2024 a Canadian tribunal ordered Air Canada to honor a bereavement refund policy that its own chatbot had invented. The airline argued that the chatbot was a separate legal entity responsible for its own statements, and the tribunal called that submission remarkable. Air Canada was ordered to pay CA$812.02. I understand that 812 dollars sounds small, but the tribunal treated a chatbot statement about a policy as the company's own statement, and that ruling applies to any company.
In April 2025 the AI support bot at Cursor invented a policy limiting logins to one device at a time, which did not exist. The answer spread on Reddit and Hacker News before anyone at Cursor saw it, the original poster on Hacker News wrote that dozens of users cancelled, and cofounder Michael Truell corrected the bot in public.
In June 2026 KPMG withdrew a report on agentic AI after GPTZero found that only 5 of its 45 citations matched their sources and that several named company case studies were misdescribed, and a misdescribed case study of that kind lands on the company named in it as much as on the publisher. In May 2026 the Canadian musician Ashley MacIsaac sued Google after an AI Overview described him as a sex offender and one of his concerts was cancelled.
Air Canada and Cursor at least ran the bot that did the damage, and KPMG published the report itself. When ChatGPT states a wrong price for you, no one at your company saw the answer and no one can reply to it, and the buyer still reads it as something you said.
Why models invent facts about your company
A model invents a fact when the sources it read contain no answer. We went through the mechanics in what do_not_infer does, and why LLMs hallucinate about your brand. The short version is that language models are built to complete patterns, and "I do not know" is a pattern they complete very rarely. When a buyer asks what your product costs and your pricing sits behind a sales call, the model reaches for what companies like yours usually charge and states that number as yours.
Hallucinations therefore concentrate where your public facts are thinnest. Claude invented a pricing tier for cloudweld.ai in our own wince test, and a pricing page that a crawler can read plain and simple is the first fix on our list.
Vendor data points the same way, with the caveat that it comes from companies selling the fix. Profound's August 2026 review of 158,000 AI claims found that pricing made up 12 percent of all claims and 24 percent of the inaccurate ones, and that 54 percent of brands had an inaccurate claim that cited the brand's own content as its source. Oumi's April 2026 test of Google AI Overviews on SimpleQA found about 91 percent of answers correct while only 39 percent were both correct and fully supported by the cited sources, so I check the citation and the claim separately.
A company with gated pricing and a blog last touched two years ago leaves the model far more to guess at than one with a public price list and current docs.
What a monitoring setup looks like
A monitoring setup has four parts: a prompt set, a schedule, a severity rule and a correction path. You can put the first version together in an afternoon, and I would start with that rough version this week instead of waiting for the polished one, since the models retrain on their own schedule and the answers about you change with them.
A prompt set
Write down the questions where a wrong answer costs you money. The first two on the list are what it costs and how it compares to the two competitors you lose deals to, and ten to twenty questions of that kind are enough to start. These are the same buyer questions we built Brand Tracker around, so if you already run it you have this list.
A schedule
A single check tells you what the models said on one day. Answers shift when a model retrains or when your own pages change, so weekly is the minimum and daily makes sense for pricing and product claims. The Tow Center figure is an average across 1,600 queries, so one clean run of your own questions says little about the next one.
A severity rule
A wrong founding year and a wrong security claim should not trigger the same response, so I would sort alerts into three tiers before the first one arrives. Minor covers a wrong founding year or an old tagline, and that goes into the next content pass. Commercial covers a wrong price or a feature you do not have, and that one I would fix at the source inside the week, since a buyer can walk away from a deal over it. Legal means a false claim about security or compliance, and that gets fixed the same day with whoever owns legal in the loop. Decide the tiers before the first alert arrives, because once it has arrived the discussion turns to how serious it is, and the wrong answer stays online for as long as that discussion takes.
A correction path
Our team plans this part first, because an alert that no one is able to act on is one more number in a report. When you find a hallucination there is no vendor form to file, and every vendor in this category, ours included, will tell you that the correction happens on the pages the model reads and that monitoring only tells you where the drift is. Which tools carry that correction inside the product and which stop at the report is the subject of which tools connect monitoring to your analytics.
Suppose Claude tells buyers your platform has no API. Start with your own site. An API documented only inside a PDF, or rendered by JavaScript that crawlers do not execute, has effectively never been seen by the model, and stating it in plain crawlable text is what we built Brand Intelligence to publish. Then look at the review sites and comparison posts the models lean on, because if G2 and two listicles omit your API, the model is repeating what G2 and those two posts say, so the fix is on their pages. After that, keep the question in your prompt set and watch for the answer to change, and once it changes after your fixes have propagated you have a reason to believe the source change did it. We covered the third party layer in more detail in why ChatGPT does not mention your company.
Who needs this beyond software companies
Any organization whose buyers or stakeholders ask an AI about it has this problem. A university has it with tuition and admissions deadlines, and a hospital has it with opening hours, and the person who gets the outdated hours shows up at a closed door.
The line between monitoring and obsessing
Obsessing over every answer is a reasonable worry, and one I share. Monitoring pays for itself only when an alert changes one of your pages. If a month of alerts has changed nothing, the prompt set is asking about things that do not affect a deal and I would cut it to the ones that do.
To be honest, if your company has one product and a public pricing page, your hallucination risk is low, and a monthly manual check with the 5 prompts may be all you need, and I say that even though we sell the category. Automated monitoring starts to make sense once your facts change often, or once nobody owns the checking and it stops happening.
FAQ
What is a brand hallucination?
A false factual claim an AI model makes about a company, such as a price you do not charge or a feature you never built, stated with the same confidence as the true facts around it.
How often do AI models hallucinate?
Handed a document to summarize, the flagship OpenAI and Anthropic models invent facts in roughly 5 to 12 percent of outputs on Vectara's hallucination leaderboard, a grounded summarization test that sets a floor for questions about your brand. Asked to identify the source of a real quote, eight AI search engines failed more than 60 percent of the time in the Tow Center's 2025 study.
Can you stop AI models from hallucinating about your brand?
You can shrink the gaps that cause it. Models infer when verified facts are missing, and publishing your pricing and your feature list in plain crawlable text leaves the model less room to fill on its own.
How do you correct a wrong AI answer?
You fix the sources. Update your own pages so the true fact is stated in plain text, correct the third party pages the models cite, and keep the question in your prompt set until the answer changes. There is no form to file with the model vendor.
Has a company ever been held responsible for an AI hallucination?
Yes. In February 2024 a Canadian tribunal ordered Air Canada to honor a bereavement policy its chatbot invented, and rejected the argument that the bot was responsible for its own statements.
Is this different from social listening or brand monitoring?
Yes. Social listening tracks what people say about you, while hallucination monitoring tracks what machines assert about you in answers your buyers treat as neutral fact.
Run your buyer questions and see which answers you would wince at
Ooky runs your prompt set on a schedule across the engines on your plan and keeps every answer, so you can see the day one changed and which page to fix. Start with the ten questions your buyers ask before the first call.
Sources
- Vectara. "Hallucination Leaderboard." HHEM-2.3, updated May 11, 2026. Retrieved 2026-08-24. https://github.com/vectara/hallucination-leaderboard
- Vectara. "DeepSeek-R1 hallucinates more than DeepSeek-V3." January 2025. Retrieved 2026-08-24. https://www.vectara.com/blog/deepseek-r1-hallucinates-more-than-deepseek-v3
- Jaźwińska, K. and Chandrasekar, A. "AI Search Has a Citation Problem." Tow Center for Digital Journalism, Columbia Journalism Review, March 6, 2025. Retrieved 2026-08-24. https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php
- Nieman Lab. "AI search engines fail to produce accurate citations in over 60% of tests, according to new Tow Center study." March 2025. Secondary coverage. Retrieved 2026-08-24. https://www.niemanlab.org/2025/03/ai-search-engines-fail-to-produce-accurate-citations-in-over-60-of-tests-according-to-new-tow-center-study/
- Civil Resolution Tribunal of British Columbia. Moffatt v. Air Canada, 2024 BCCRT 149. Decided February 14, 2024. Retrieved 2026-08-24. https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html
- Garcia, M. "What Air Canada Lost In 'Remarkable' Lying AI Chatbot Case." Forbes, February 19, 2024. Retrieved 2026-08-24. https://www.forbes.com/sites/marisagarcia/2024/02/19/what-air-canada-lost-in-remarkable-lying-ai-chatbot-case/
- Edwards, B. "Company apologizes after AI support agent invents policy that causes user uproar." Ars Technica, April 21, 2025. Retrieved 2026-08-24. https://arstechnica.com/ai/2025/04/cursor-ai-support-bot-invents-fake-policy-and-triggers-user-uproar/
- Hacker News. Discussion thread on the Cursor support bot device-limit policy. April 2025. Retrieved 2026-08-24. https://news.ycombinator.com/item?id=43683012
- Patterson, M. "Air Canada's Chatbot Walked So Cursor's Chatbot Could Ruin." Help Scout, July 30, 2025. Retrieved 2026-08-24. https://www.helpscout.com/blog/ai-curse-of-cursor/
- The Register. Report on KPMG withdrawing its agentic AI report after GPTZero's citation check. June 12, 2026. Retrieved 2026-08-24. https://www.theregister.com/ai-and-ml/2026/06/12/kpmgs-ai-report-turns-into-a-demo-of-ai-hallucinations/5255029
- TechCrunch. Report on KPMG pulling its AI usage report over apparent hallucinations. June 13, 2026. Retrieved 2026-08-24. https://techcrunch.com/2026/06/13/kpmg-pulls-report-on-ai-usage-due-to-apparent-hallucinations/
- Global News. Report on Ashley MacIsaac's defamation claim against Google over an AI Overview. May 2026. Retrieved 2026-08-24. https://globalnews.ca/news/11829836/alleged-defamation-ashley-macisaac-google/
- Profound. "Where do inaccurate AI claims come from?" August 2026. Vendor data. Retrieved 2026-08-24. https://www.tryprofound.com/blog/where-do-inaccurate-ai-claims-come-from
- Oumi. Study of Google AI Overviews against SimpleQA and their cited sources. April 2026. Retrieved 2026-08-24. https://oumi.ai/blog/oumis-study-finds-50-of-ai-overviews