Why AI Gets Your Brand Wrong: It's Usually Not Hallucination — Old Titles, Retired Product Names and Stale PDFs Are Poisoning AI Answers
Hi everyone, this is Neo.
Let me run a thought experiment with you. You run a cross-border ecommerce brand. Three years ago your company got acquired by a parent group, the CEO left, and the business is now led by a President & General Manager. So you ask ChatGPT: “Who is the CEO of [Brand]?” It answers with your former CEO — someone who left four years ago. Annoyed, you go check the website: the new leadership info is right there, clearly written. Why can’t the AI see it?
Here’s the twist that might surprise you: the AI didn’t miss the new information. Your website just has too many “old versions” lying around — and the old versions are winning the retrieval race.
This isn’t an isolated case. Carolyn Shelby of search consultancy CSHEL, a Search Engine Journal columnist, put the phenomenon on the table in her September 7 article — tellingly titled “Your Biggest AI Search Risk Is Conflicting Information About Your Brand.” Her core point cuts deep:
When the information about a brand is outdated or wrong, the problem is often not too little data. It’s too much data. The brand already has too many versions of the truth online — and the version that best matches the user’s question isn’t the current one.
1. Why the old version keeps winning
Before anything else, understand the mechanism — otherwise you’ll keep treating the wrong disease.
In the traditional search era, Google could rank the page that calls you “Company A” and the page that calls you “Company B” at the same time, and let the user figure out which one is current. The engine never made the call for you.
In the AI search era, that changed. ChatGPT, AI Overviews, Perplexity and their peers retrieve a set of sources and assemble them into a single answer. Here’s the catch: even when the underlying evidence contradicts itself, the AI’s answer always looks settled. It won’t add a footnote saying “two conflicting versions of this company’s facts exist, please verify.”
Shelby gives a painfully real example. When you ask “Who is the CEO of [Company]?”, the question itself carries an outdated assumption — that the company still has a CEO. The AI searches around “company name + CEO,” which naturally favors pages containing that exact relationship: old press releases, old bio pages, interviews from years ago, acquisition announcements. Meanwhile the company’s current leadership page uses new vocabulary — “SVP,” “brand president,” “GM of the business group” — that doesn’t map to your question.
An LLM cannot cite a source it never retrieved. If the content describing the new structure never connects itself to the old terminology, it never enters the retrieval set — so the old answer wins, because only it matches the language of the question.
Read that twice. Your new page isn’t failing on authority — it’s failing because it doesn’t sit on the vocabulary track of what people actually ask.
And don’t underestimate how much “information debt” a typical website carries. I’ve asked quite a few independent ecommerce sellers to run a quick self-audit, and almost everyone digs up a pile of these:
- Old pages from a past site redesign still living on subdirectories or old subdomains
- Media kits, sales decks and PDF whitepapers from years ago still being indexed by Google
- Product renames/revamps where the old product-name pages were never 301-redirected, so two generations of pages coexist
- Discontinued plans, expired certifications, and dropped service regions still sitting in the help center
- Product documentation using language the marketing team abandoned two years ago
- Executive bios whose opening paragraph still carries an obsolete “Founder & CEO” title
- App store listings, job posts, speaker bios and partner pages each telling a slightly different story
Some of this content was accurate when published and is now stale; some was contradictory from day one. It didn’t matter before — Google ranked it all. It matters now: the AI picks whichever version best matches the question’s wording and presents it as the one truth.
2. Why publishing a “new page” doesn’t fix it
Most brands’ first instinct: publish one more official page / a new site / another FAQ with the correct facts, and the problem solves itself.
Shelby’s answer: publishing is not the same as correcting. Accurate information written in the wrong vocabulary stays invisible to the question people actually ask. If you write “Jane Smith is SVP and GM,” but the user asks “who is the CEO,” nothing bridges those two phrasings, so retrieval never connects them.
The fix has a name: bridge content — sentences written specifically to connect the terminology people still use with the reality that replaced it. Here’s the template she offers:
“Following the acquisition, [Company] no longer has a standalone CEO. Jane Smith now leads [Company] as SVP and general manager within [Parent Company Group].”
Why this sentence works: it contains the old term people still search for (CEO), states the relationship explicitly (no more CEO → now SVP/GM), so retrieval can match “CEO” queries without wrongly assigning the title to anyone.
The same logic applies everywhere:
- Product renamed: “We used to be called A; we’re now B, same features and pricing.”
- Plan retired: “Our old Basic plan ended in June 2026; the Starter plan replaces it.”
- Merger: “[Brand] has been acquired by [Group] and now operates as a [Group] brand.”
- Service area change: “We no longer serve the EU; EU customers, please see [Brand EU].”
One principle: never assume users know the new name. Write down the relationship between the words they still use and the reality that took over.
Shelby also flags a trap people fall into constantly: once you find an incorrect AI answer, the right move is to trace it back through the citations and evidence chain — not to bury it under a pile of freshly published pages. Handle each link in the chain differently:
- Owned pages: update what can be updated (current bios shouldn’t keep obsolete titles in the opening copy), consolidate or 301 what should be merged. Old announcements shouldn’t be rewritten to pretend history happened differently — but they can carry a clear date, a “status: outdated” note, and a link to the current authoritative page. PDFs and stale sales materials should be retired or labeled archival.
- Third-party pages (media coverage, directories, partner pages): you can’t force a publication to rewrite an accurate five-year-old story. What you can do is make the canonical current explanation extremely easy to find, and politely ask directories and partners whose job is to stay current to update their records.
3. Upgrade your content inventory into a “brand claim audit”
A traditional content inventory records URLs, traffic, rankings, and maybe conversions. Shelby argues that in the AI-search era you need a new instrument: a brand claim audit — one that records what factual assertions those assets actually make, and what language people are likely to use when they ask about those facts.
For every important claim, you need three answers:
- How many versions of this fact are currently live on the web? (Where’s the old one, where’s the new one?)
- What vocabulary does each version use? (“CEO” vs. “SVP”? “Basic plan” vs. “Starter plan”?)
- What wording is the user most likely to search with? (That decides where you need to build the bridge.)
And the audit scope must not stop at HTML pages. Shelby calls out a specific list of blind spots:
Media kits, downloadable sales materials, help-center content, schema values, product feeds, app-store listings, speaker biographies, job listings, old subdomains — and ask sales, support, HR, product, legal and marketing which materials they publish without the web team ever knowing.
Most traditional SEOs have never audited any of that. But an AI’s retrieval scope is the whole web — the PDF your sales team sent a client three years ago may be easier for an AI to cite than your own homepage.
4. Stop tracking “was I mentioned” — evaluate the answer itself
Most “AI visibility reports” on the market today (including paid tools) revolve around one metric: how many prompts mention or cite your brand. Shelby warns that this kind of presence creates false confidence.
A brand mention is not a win if the answer names a former executive, quotes an old price, attributes a discontinued feature to the current product, or misunderstands the relationship between the brand and its parent company.
The right approach: for high-value prompts, evaluate the answer itself. Is it accurate? Is it current? Does it answer the user’s real intent — or does it just accept the outdated false premise embedded in the question? When the answer is wrong, inspect the citations and reproduce the likely searches before assuming the model hallucinated. Most of the time, it simply retrieved one of your old versions.
This isn’t just Shelby’s opinion. Google’s own documentation on generative AI in Search states that these features use retrieval-augmented generation and query fan-out to pull and synthesize information from the Search index — garbage in, garbage out. Third-party data is equally damning: a 2026 arXiv measurement study of Google AI Overviews analyzed 55,393 trending queries and found that 11.0% of decomposed atomic claims were not supported by the pages cited. Errors with citations attached — and a huge share of those trace back to your own stale content.
Brand information governance, in other words, is reverse SEO: instead of getting AI to find more about you, you’re now engineering it to find only the correct version.
5. An action checklist for independent site owners
Enough theory — here’s a workflow you can start today.
Step 1: Build a high-risk question list. First decide which facts affect purchase decisions — price, plan names, feature availability, supported regions, integrations, compliance, certifications, product names, founders/executives. Ask in buyer language (not brand language): “How much does [Brand] cost?” “Does [Brand] still support PayPal?” “Who runs [Brand]?” Remember: one clean answer proves nothing — generative answers vary, so run the same prompt repeatedly and track the error rate.
Step 2: Ten-minute content-debt audit. Run two kinds of searches:
site:yourdomain.com— reverse-chronological scan for live old titles and retired product terms- Exact-phrase searches (in quotes) for your dead product names, retired plan names and old taglines — see which third-party pages still repeat them
Step 3: Prioritize by business risk, not by page age. Facts that affect a buying decision (price, features, stock) come first; cosmetic facts (bio wording) second. Follow the principles above: update what you can, date-stamp old announcements, put bridge content on the “old term → new reality” path, and delete or archive stale PDFs.
Step 4: Close the verification loop. After fixing, go back to the high-risk questions and re-run the same prompts to confirm the answers actually changed. Fix-without-verify is fix-without-effect — AI caches and index refreshes take time, so a follow-up check a week later is the real proof.
Neo’s take
This article deserves a careful read by every founder doing cross-border ecommerce, because it kills a widespread misconception: that “the AI got my brand wrong” equals “hallucination” equals “nothing I can do.” The truth: most wrong AI answers can be traced, evidence-in-hand, back to your own website — that pile of old pages you assumed nobody was looking at.
Put this together with what we’ve covered before and the picture completes itself:
- We previously covered Duane Forrester’s “information vacuum” problem: when AI has nothing on your brand, it describes you as your competitor or an industry average — the disease of too little information.
- Today’s Shelby piece is the mirror image: too much information that contradicts itself — your old versions crowd the current one out of the retrieval set.
Two diseases, one prescription: brand fact governance. A vacuum needs authoritative content poured in; a conflict needs a single version of the truth enforced. This is also why so many big brands burn money on GEO with nothing to show for it — they’re busy publishing fresh content and chasing citations while their own back yard of contradictory legacy pages quietly feeds the AI wrong answers. Here’s a KPI I’d add to every content team’s dashboard for the AI era: for every critical fact the company publishes, only one “current version” may exist on any public channel.
One more thing for Chinese cross-border teams specifically: your content debt is usually heavier than a Western brand’s — overseas expansion almost always involves rebrands, multi-platform storefronts, and a parade of agencies, each leaving behind its own version of the story. Spend one afternoon running a brand claim audit. It might be the highest-ROI SEO work you do all year.
I’ll leave you with the best line in Shelby’s piece:
Sometimes the strongest content strategy is fixing the old content you forgot you published.