When ChatGPT Gets Your Company Wrong: How to Find and Fix Stale AI Facts (2026)
When an AI engine describes your company wrongly, it is one of four separate problems. Diagnose which one you have before you touch anything, because three of the four fixes will do nothing.
On this page
- The 30-second answer
- The four failure modes, and what each one actually needs
- Why does an AI engine get your company wrong at all?
- How do you find out what AI is saying about you, without fooling yourself?
- What can you actually make each engine do?
- How long does a fix take?
- Does schema markup fix this? Does blocking the bots?
- When does a wrong AI answer become a legal problem?
- What we would actually do, in order
- Frequently Asked Questions
The 30-second answer
When an AI engine describes your company wrongly, it is one of four different problems, and each one takes a different fix. A hallucination has no source to correct. A stale fact has a source that needs updating and recrawling. A mis-attributed fact is sitting on the wrong company's name. Contradicting sources means both the old and the new version of you are live and both get retrieved. Diagnose which one you have before you touch anything. Each of the four fixes is inert against the other three modes, so the right repair aimed at the wrong problem buys you nothing.
The four failure modes, and what each one actually needs
| What went wrong | Why it happens | The fix lever | Realistic timeline |
|---|---|---|---|
| Hallucination. The answer states something no source says, sometimes citing a page that does not exist. | The model generated plausible text with nothing behind it. | Report it through the vendor's channel, then publish an authoritative page that answers the question properly. | No vendor publishes one. Retrieval-grounded surfaces can shift in days. Model memory does not shift until a retrain. |
| Stale fact. The answer was true eighteen months ago. | Temporal misalignment. The model reflects its data-collection window rather than the present. | Update the source page, then get it recrawled. | Not published. The retrieval path can flip quickly. Parametric memory does not. |
| Mis-attribution. A real fact, attached to the wrong company. | Entity-based knowledge conflict. Models over-rely on memorised entity associations. | Strengthen entity disambiguation, and fix the third-party pages that make the confusion look correct. | Not published. Our read rather than a measured figure: the slowest of the four. |
| Contradicting sources. Old you and new you are both live, and both get retrieved. | Inter-context conflict: contradictory evidence among the retrieved passages themselves. | Retire, redirect or correct the contradicting pages. | Not published. Our read rather than a measured figure: the most tractable of the four. |
A hallucination has a published definition: output that "cannot be verified from the source content." Reporting it is half the job. The authoritative page is the other half, so the retrieval layer has something real to reach for.
Stale facts are the data-collection window showing through. Anthropic publishes Claude Haiku 4.5 with a training cutoff of Jul 2025 but a reliable knowledge cutoff of Feb 2025. Updating the page is only the start. Allow the search-surfacing crawlers so the corrected page can be picked up, or the correction never travels.
With mis-attribution, the wrong pairing is stubborn, and we rank it slowest because you are arguing with the model's memory rather than with a page. Contradicting sources are the common case after a rebrand, a pivot or a pricing change. That one is the most tractable, and the one you can genuinely close. Citations concentrate, so the list of hosts that matter is short.
Two of those four timings are our ranking rather than anyone's measurement. No published study measures how often each mode occurs in the wild, or how often each fix works. The taxonomy is in the peer-reviewed literature. The triage statistics are not.
Most companies skip the diagnosis and go straight to the fix they already know how to do. That is why so much of this work produces nothing.
Why does an AI engine get your company wrong at all?
There are two separate knowledge stores behind every answer, and they fail differently.
The original retrieval-augmented generation paper set out the split: models combine "pre-trained parametric and non-parametric memory for language generation." Parametric memory is the model's weights. Non-parametric memory is a document index it can retrieve from at answer time. The same 2020 abstract already named the problem you are dealing with today, listing "providing provenance for their decisions and updating their world knowledge" as open research problems. Six years later they are still open.
The vendors document the split themselves, through their crawlers. OpenAI runs GPTBot to make its "generative AI foundation models more useful and safe," and OAI-SearchBot to "surface websites in search results in ChatGPT's search features." Perplexity is blunter: PerplexityBot exists to "surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models." Anthropic splits the same three ways.
This matters because it tells you which lever to pull. Google states that a page becomes eligible for AI Overviews and AI Mode by ordinary means, that "a page must be indexed and eligible to be shown in Google Search with a snippet," and that "There are no additional technical requirements." Its AI surfaces then use what Google calls a query fan-out technique, "issuing multiple related searches across subtopics and data sources." We wrote up what fan-out means for the prompts you track separately, because it changes what you are actually optimising for.
The uncomfortable part: most of what an engine says about you is grounded in pages you do not own. One study of 167,551 grounded citations across Perplexity Sonar Pro, Gemini 3.1 Pro and GPT-5.4 found 85.7% of citations pointing to non-brand domains and 14.3% to the brand's own site. Even the most self-referential brand in that corpus, Tatra Banka, drew only 34.4% of its citations from its own pages. A separate study of 100,000-plus prompt responses reports close to the inverse, roughly 78% of citations going to corporate websites. The direction both support is the one that matters: you do not control most of the evidence.
One caveat before the numbers start stacking up, and it covers every measurement figure quoted here rather than only the two above. Almost all of the published measurement in this field comes from a handful of 2026 preprints whose authors work for AI-visibility vendors, and none of it has been replicated by anyone without a product to sell. That includes the 85.7 and 14.3 split, the Tatra Banka figure, the 78% counter-reading, the domain-concentration numbers, the Wikipedia share, the appearance-rate bands, the sampling-design numbers and the sentiment-volatility multiple. They are the best available readings of direction and order of magnitude. They are not constants, and a research programme this young should not be quoted like one.
How do you find out what AI is saying about you, without fooling yourself?
This is where good intentions produce garbage data. The instinct is to open ChatGPT, ask about your company ten times, and draw a conclusion.
A variance-components study across 12,933 responses measured what actually moves the reading. Query language accounted for 26.5% of variance. Model identity accounted for 1.6%. Brand identity accounted for 1.5%. The allocation rule the authors landed on is the opposite of standard practice: "a repeat past the fifth reduces relative-error variance by about 0.0003," while "language diversity reduces relative-error variance about fifteen times as much as five more repeats." Their conclusion: "reliability is bought by breadth across languages and models, not by depth of repetition."
A second team, sizing runs for a different measurement, found the standard error of per-brand detection rate dropping below 0.10 at seven runs. Their framing is the right one to adopt: treat visibility "as a distribution rather than a single-point outcome."
Two more things will distort your reading.
Your company's size sets your baseline. Measured appearance rates across 100,000-plus prompt responses: global household names 73%, established mid-market brands 44%, niche and small brands 11%. If you sit in the 11% band, absence and error look identical until you sample properly.
And being well known raises the odds of fabrication. A study of 100 Hungarian B2B entities found Tier 1 brands producing 52.69% fabricated citations against 37.87% for Tier 3, with regulatory-framed queries pushing fabrication to 56.77%. The authors call it the Brand Hallucination Paradox, and explain it as familiarity creating "stronger surfaces for plausible but incorrect completions." One language, one market, so do not carry those percentages into your own deck. Carry the mechanism.
For direct observation rather than sampling, read your own logs. Server logs are the highest-fidelity signal we have for what AI bots actually fetch, because they are observation instead of estimation. Cloudflare measured AI bots at an average 4.2% of HTML requests across 2025, ranging from 2.4% to 6.4%. And Bing Webmaster Tools has published a free AI Performance report since February 2026, covering Microsoft Copilot and AI-generated summaries in Bing. Grounding queries, the phrases the AI used when it retrieved your page, were in that first release. A June update added Intents, Topics, Citation Share and period comparison. Microsoft bounds the whole thing honestly: it "does not indicate ranking, authority, or the role of any page within an individual answer."
What can you actually make each engine do?
Here is every documented correction channel we could find, and what it really promises.
| Engine | Documented channel | What it commits to |
|---|---|---|
| Google AI Overviews | Thumbs up or down, "Report a problem," free text | Nothing. Google's own line is "AI Overviews can and will make mistakes." |
| Google Knowledge Panel | Claim the panel, then per-fact feedback with verification links | Reviewed within a few days, criteria published per field, but Google will not write you a new description. |
| Google Search index | Removals tool, for pages you host | Removal "within a day," lasting "about 6 months." Pages you own only. |
| ChatGPT | Privacy Portal, personal-data removal, with ID | A case-by-case review of personal data about a person. |
| Perplexity | Flag icon under the answer, support ticket, or email | Nothing published. Its own issue list does name Misinformation and Outdated information. |
| Bing and Copilot | Block URLs, 404/410, NOINDEX, IndexNow, plus a tool for non-owners when results are stale | Blocks up to 90 days, most processed under 12 hours. |
| Anthropic Claude | robots.txt blocking of ClaudeBot, thumbs-down button, feedback address | Future crawling only. Nothing about correcting what a model already holds. |
Google's Knowledge Panel is the most any of them commits to, and the only published review window here: verified feedback reviewed "within a few days" with an emailed resolution update, though Google adds that "This can sometimes take more time." Criteria are stated per field. A title change needs "substantial evidence that our automated systems didn't make the most representative selection"; a description needs "strong evidence" plus proof you asked the source first. Google can delete an unsupported description but "can't create a custom description."
OpenAI's portal handles personal data about a person, one case at a time. Removal is scoped to ChatGPT and does not touch external sites or search engines. Bing adds its own caveat: search engines cannot delete content from a website, so go to the site owner. Anthropic's documented route is robots.txt blocking of ClaudeBot, plus the in-product thumbs-down button and a feedback address named on its own incorrect-responses help page.
Not one of those channels commits to correcting a factual claim inside an AI answer. They are suggestion boxes and index-hygiene tools. Google does run one accuracy-gated business-fact channel, and it is not in the table because it never touches an AI answer: Business Profile edits are reviewed against Google's business-information guidelines, usually inside ten minutes but sometimes taking up to 30 days, and Google might not approve a change when "it can't confirm its accuracy." That is the only place in the stack where somebody checks whether a business fact is true before publishing it, and its scope stops at your opening hours and your address. The Knowledge Panel gets closest to an entity fact, and its ceiling is deleting a description it agrees is unsupported.
Which is why the real work is upstream. Citations are concentrated: in that 167,551-citation corpus, "80% of citations come from about 18% of domains," and "Half of all citations come from just 547 hosts." That concentration is the good news. You are not fixing the web. You are fixing a short list.
One correction to a common belief before you build a plan around it. Wikipedia measured at 3.9% of grounded citations in that corpus, across 128 brands in 12 European markets, counting only citations the engines actually grounded an answer in. The 22% and 47.9% figures circulating on agency blogs do not survive the one corpus study that measures citation share directly. Wikipedia led as a single host in 11 of 12 languages, so it matters, but it is not the lever people think. And most B2B SaaS companies cannot pull it anyway: Wikipedia requires "significant coverage in multiple reliable secondary sources that are independent of the subject," and on independence it is unambiguous, "Only unpaid sources count." Press releases, marketing material and sponsored posts are all explicitly rejected.
How long does a fix take?
Nobody knows, and anyone who gives you a confident number is guessing.
We looked for a published figure on how long after fixing a source an AI answer changes. No vendor publishes one. Google's crawler documentation covers crawl rate and says nothing at all about how fast a change gets reflected. The Knowledge Panel is the one place Google does put a review window in writing, a few days for verified feedback and sometimes longer, and that window measures how long until somebody reads your suggestion. The timelines that do exist, Google's one-day removal and Bing's under-twelve-hours block, are about whether a page is indexed. Neither commits to when an answer stops repeating a wrong fact. The specific numbers you will see quoted, a two-day Perplexity median or a 3.4-week citation half-life, all trace back to tool vendors with no inspectable method.
The mechanism is more useful than a fake number, and there is a documented natural experiment. In the Norwegian complaint noyb filed against OpenAI, ChatGPT stopped repeating a fabricated claim once it began searching the web for information about the complainant. But noyb also recorded that "the incorrect data may still remain part of the LLM's dataset," with no certainty it can be erased "unless the entire AI model is retrained."
So: the retrieval layer can flip fast. The model's memory does not flip at all until a retrain. That is the honest answer, and it sets expectations correctly.
Our own reading of citation speed matches it. We track three separate speeds, and conflating them is how agencies oversell. First pickup can happen in hours to a day on a low-competition prompt. Holding a slot in the cited-source set takes weeks. Dominant share of answer on a competitive cluster takes months. We have written up the three speeds in full, because clients are routinely sold the fast one and billed for the slow one.
Worth knowing which surface moves first. Google AI Overviews sits on the live index and updates within hours. Perplexity rebuilds on a daily to weekly cycle. ChatGPT's base training is fixed at a cutoff and refreshes on major model releases, so recent pages go missing from its answers even with retrieval. On Toku, Google AI Overviews accounted for 57% of total AI mentions. If you are checking ChatGPT first because it is the tool you use yourself, you are checking the slowest panel first.
Does schema markup fix this? Does blocking the bots?
Two popular answers, both wrong in instructive ways.
Schema markup. Google documents Organization markup and sameAs for entity disambiguation, saying it helps Google "disambiguate your organization in search results," with properties like iso6523 and naics working "behind the scenes." That is real and worth doing. But for AI Overviews and AI Mode specifically, the same vendor closes the door: "You don't need to create new machine readable files, AI text files, or markup to appear in these features," and "There's also no special schema.org structured data that you need to add." Anyone telling you sameAs makes AI answers more accurate is stating practitioner belief rather than documented behaviour. Ramp's own A/B test serving formats to AI crawlers found plain markdown outperforming injected structured data, with a caveat the team recorded itself: the markdown variant was served to a broader set of bots than the other two, so some of the gap may be targeting rather than format. Hold it loosely. The Google documentation above already carries the argument on ground the vendor signed.
Blocking crawlers. This one is actively counterproductive, and it is being done at scale. Blocking AI bots does not protect your accuracy. It removes the path a correction travels. Both vendors running a search-surfacing crawler tell you to allow it. OpenAI: "To help ensure your site appears in search results, we recommend allowing OAI-SearchBot in your site's robots.txt file." Perplexity says the same about PerplexityBot. Cloudflare, meanwhile, found AI crawlers "were the most frequently fully disallowed user agents found in robots.txt files." Companies are blocking the bot that would have picked up their corrected page.
Note the asymmetry, and do not flatten it. robots.txt is a reliable control for automated training and indexing crawlers. It is unreliable for the live per-question fetch: OpenAI says of ChatGPT-User only that "robots.txt rules may not apply," and Perplexity says its equivalent "generally ignores robots.txt rules." Neither claims an exemption. Both hedge. The rule works per user agent rather than per vendor.
And llms.txt is not the answer either. Across 137,210 domains in May 2026, 97% of published llms.txt files received zero requests, and AI retrieval bots accounted for just 1.1% of the requests that did arrive. The study's verdict: "If your goal is showing up in ChatGPT, Perplexity, or AI Overviews, an llms.txt file is largely decoration."
When does a wrong AI answer become a legal problem?
The picture changed in July 2026, and most marketing teams have not caught up.
In Starbuck v. Google LLC (Delaware Superior Court, C.A. No. N25C-10-211 MAA, decided 2026-07-24), Google's motion to dismiss was denied in full. The opinion's own words: "This Opinion DENIES Google's motion in its entirety." The court called it "a new frontier for defamation law" while resolving the motion "based on established defamation caselaw."
Three details are directly useful to anyone documenting a wrong answer.
Written notice to the vendor's legal department mattered. The plaintiff first escalated through public posts tagging Google executives, which went nowhere. Then, "on July 31, 2025 and August 12, 2025, Starbuck sent written correspondence to Google which was received by Google's legal department." The court held it "possible the Legal Department Notices put the proper Google personnel on notice of the falsity of the Outputs," and that failing to rectify afterwards was enough to reach discovery on actual malice.
Fabricated sources are evidence in themselves. The court noted that the plaintiff "alleges Google AI fabricates sources, which courts have determined can support a finding of actual malice." Practically: screenshot the cited sources as well as the claim. An answer citing a page that does not exist is a different animal from an answer citing a real page that is out of date.
A second live case makes that same point with a company in the plaintiff's chair, which is the seat most readers of this article occupy. LTL LED, LLC, trading as Wolf River Electric, is a Minnesota solar installer suing Google over an AI Overview that said the company faced a Minnesota Attorney General lawsuit over deceptive solar sales. It was never a defendant in that action. Its initial disclosures put damages between roughly $110 million and $210 million, and it pleaded one customer terminating a $150,000 contract after reading the answer, despite the CEO telling that customer the claim was false. The complaint's central allegation, as the Volokh Conspiracy's write-up of the remand decision renders it, is that Google "cited numerous sources in support of its false assertions; however, none of the referenced materials in fact contained the information Google claimed they did." Fabricated citations, with an invoice attached. The matter is still running: a federal judge sent it back to Minnesota state court in January 2026 because Google filed its removal notice late, and Volokh's read is that the plaintiff "appears not to be a public figure, and appears to have evidence of tangible economic losses. That makes its case considerably stronger."
Disclaimers did not end it, and Google ran the argument on two elements rather than one. Google's position was that its accuracy warnings foreclose any reasonable third party from relying on the outputs, which would defeat the element of publication. The court declined to decide, because "The disclaimers Google references are not identified in the Complaint or attached as an exhibit." Later in the same opinion, under actual malice, Google ran the warnings point again, arguing that this "negate[s] an inference of malice." The court deferred that branch on the same evidentiary ground, "the scope of Google's purported disclaimer was not alleged in the Complaint and has not been provided as an exhibit," and on three further ones. Relying on the disclaimer was "inappropriate at this stage of the proceedings." Deferred, not accepted.
The decision usually cited against all of this actually reinforces it. In Walters v. OpenAI, a Georgia court granted OpenAI summary judgment in May 2025 on what the reporting of the order describes as three independent grounds, and disclaimers were a supporting factor in two of them, decisive in neither. On the first ground, no defamatory meaning, warning language "weighs in the determination" of how a reasonable reader would read the output rather than settling it. What settled it were the facts of the exchange: ChatGPT told the requesting user it could not open the link he had pasted, the user was holding the real complaint, he established "within about an hour and a half" that the output was untrue, and he testified that he "understood that the machine completely fantasized this." On the second ground the court found no negligence, with OpenAI's warnings to users counted as further support for that finding, and separately treated Walters as a public figure who produced no evidence that anyone at OpenAI knew the output "would probably be false." On the third, he could not show actual damages, and he lost punitive damages outright because Georgia law makes a libel plaintiff request a correction or retraction before filing and he never made one.
Sit with that last one. The plaintiff forfeited a damages category for skipping the step this piece ends on: asking the vendor, in writing, to fix the answer. Starbuck was helped by written notice; Walters was penalised for its absence. Both point the same direction, so the two decisions are consistent rather than split, and the disclaimer question is not the fault line people report. Starbuck refused to reach it for want of evidence and deferred it until after discovery. The Starbuck court names a separator that is doctrinal rather than procedural: Starbuck alleges Google AI was "deliberate[ly] engineered" to defame him, and Google "offers no caselaw addressing how a purported disclaimer as to the veracity of an AI output interacts with a claim of intentional false representation." Procedural stage and evidentiary record are two more. The Starbuck court said as much when it distinguished Walters as a case "decided on summary judgment," dismissed only after discovery had confirmed what the disclaimers said and who saw them. Eugene Volokh, who has tracked these cases from the first filing, reads Walters as tied closely to its facts and notes that a plaintiff who had alerted the defendant and been ignored might well have come out differently. That describes Starbuck.
Two honest limits. The corporate case law is young rather than absent. Wolf River is the only company-plaintiff matter we located, and it has not been tried, so nobody knows what a business can actually recover. And the routes people assume exist do not. GDPR rectification protects personal data about a natural person, so it helps a named founder more than a limited company. The EU AI Act's Article 50, applying from 2 August 2026, covers disclosure and labelling of AI content and contains nothing requiring anyone to correct an output. The FTC's Operation AI Comply targets deceptive marketing of AI products and tools built to manufacture fake content, and none of its actions gives a business a route to correct an AI answer about itself.
What we would actually do, in order
Diagnose first. Establish which of the four modes you have (hallucination, stale fact, mis-attribution, contradicting sources) before touching a page.
Sample properly. Go for breadth across engines and phrasings rather than repetition. Five runs is roughly the point of diminishing returns on repeats.
Capture the cited sources every time, alongside the wrong sentence.
Fix the short list of hosts that actually carry your citations, starting with the contradicting ones.
Let the search-surfacing crawlers in, and check your logs to confirm they came.
File through the vendor channel, in writing, and keep the record.
We run this on ourselves and it is not flattering. Our own visibility across tracked prompts is 11.33% at an average cited position of 2.65, with a sentiment reading of 60 against a 65 to 85 norm, and our weakest topic sits at 53. Sentiment is the noisier signal by a wide margin, flipping "about 6.7 times more often than whether it is mentioned at all," so a single soft reading is not a verdict. It is still ours, and it is why we treat description as a separate problem from citation rather than assuming presence solves it.
If you want the description itself to improve, the evidence points at specifics rather than assertion, though the evidence is thinner than the numbers make it sound. The one study that measures this is a simulated recommendation bake-off: three small models, a single product category, one real brand against fictional challengers. Inside that setup, product parameters explained 82.4% of ranking variance while brand identity alone explained 1.2%. The authors replicated the effect in two further product categories and warn that specification-rich domains may dilute it. So take it as directional rather than settled: named, checkable specifics look like they carry more weight in how a model describes you than adjectives about your brand do.
If you would rather not build the measurement layer yourself, that is what our GEO and AEO work is. Our free AI visibility audit checks what ChatGPT, Claude, Gemini and Perplexity currently say about your brand, and returns it in a few minutes.
Frequently asked questions
Answers to the questions readers ask most about this topic.
Can I make ChatGPT delete something wrong about my company?
Not as a company fact. OpenAI's documented route is a privacy request for personal data about a person, assessed case by case, and it is scoped to ChatGPT only. OpenAI's own page states that removing data from ChatGPT does not remove it from external websites or search engines. There is no business-facts correction desk.
How long until AI stops repeating the wrong information?
No vendor publishes a figure, and any specific number you see traces back to a tool vendor without a stated method. The mechanism is the useful part: retrieval-grounded surfaces such as Google AI Overviews can change within hours to days once the underlying source changes, while a model's internal memory does not change until it is retrained.
Will adding schema markup make AI describe my company correctly?
It helps Google disambiguate your organisation in Search and the knowledge panel, which Google documents. It is explicitly not a requirement for AI Overviews or AI Mode, where Google states no special structured data is needed. Treat schema as entity hygiene rather than as an AI accuracy fix.
Should I block AI crawlers to stop them getting things wrong?
No. Blocking removes the route a correction travels. OpenAI and Perplexity both recommend allowing their search-surfacing crawlers so your pages appear in results. You can allow the search crawler while disallowing the training crawler, which is the split OpenAI documents.
How many times should I test a prompt to trust the answer?
Around five repeats is where extra repeats stop paying, so after that, add more engines and more phrasings instead of running the same wording again, and add languages if you sell across markets. One variance study found language diversity buying roughly fifteen times more reliability than five more repeats of an identical prompt. A separate run-sizing analysis put seven runs as the point where per-brand detection settles down.
Does Wikipedia control what AI says about my company?
Less than people assume. In a study of 167,551 grounded citations across 128 brands in 12 European markets, Wikipedia accounted for 3.9%, though it was the leading single host in 11 of 12 languages. Most B2B SaaS companies also fail its notability bar, which requires significant independent coverage and counts only unpaid sources.
Is a wrong AI answer about my business defamation?
It may be, and the case law is young. A Delaware court let a defamation claim against Google proceed in July 2026, holding that written notice to the vendor's legal department could establish awareness of falsity. A Georgia court granted OpenAI summary judgment in 2025 on three separate grounds; inside the damages ground, the plaintiff forfeited punitive damages by never requesting a correction before suing. Both point the same way on putting your complaint in writing. A Minnesota company's suit over an AI Overview is still live. Talk to counsel rather than to a marketing agency about this one.



