Google's Research Just Confirmed It: AI Remembers 95% of Facts But Forgets Them When You Flip the Subject and Object — What It Means for Your Entity SEO
Hi everyone, this is Neo.
Here’s a question: when an AI fails to answer something, is it because it never learned it, or because it can’t recall it?
We’ve all assumed the former — not enough data, right? A Google research paper published in mid-August just flipped that assumption on its head.
The paper has a great title: “Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality.”
The headline finding, in one sentence: frontier models like GPT-5 and Gemini-3 encode 95-98% of facts, yet fail to directly recall 26-34% of them. And buried in there is a discovery that hits SEO like a bomb — flip the subject/object order of a fact, and the AI’s ability to recall it drops noticeably.
Let me break down what the research actually says, and what it means for how you write content on your independent site.
The framework: encoding, recall, recognition
The paper introduces a “knowledge profiling” framework that separates a model’s knowledge into three levels:
- Encoding: has the fact been stored in the model’s parameters?
- Recall: can the model retrieve it without external cues?
- Recognition: can the model pick the right fact out of a list of alternatives?
To test this, the researchers built WikiProfile, a benchmark of 2,150 real facts extracted from Wikipedia, each paired with 10 tasks (2 encoding, 4 knowledge evaluation, 4 multiple-choice recognition). They ran roughly 4.5 million responses across 13 LLMs.
The result is stark:
“Encoding is saturated; recall is not.”
Gemini-3-Pro and GPT-5 encode 95-98% of facts, yet still fail to directly recall 26-34% of them. Even with thinking enabled, they still fail on 11-12%. Recall failures account for more than 70% of GPT-5.2’s errors.
In plain language: the bottleneck for frontier models is no longer missing knowledge — it’s knowledge that’s stored but inaccessible. To use the paper’s own metaphor: the shelves aren’t empty. The keys are lost.
The finding that matters for SEO: subject/object order
The most talked-about discovery is a fresh take on the reversal curse.
Here’s the example from Google’s explainer:
“Oasis played their first gig at the Boardwalk club.”
In this sentence, the subject entity is Oasis and the object entity is the Boardwalk club — the subject appears first in the source text, the object appears subsequently.
That defines two question types:
- Direct question: the answer sits in the object position. “Where did Oasis play their first gig?” → The Boardwalk club.
- Reverse question: the answer sits in the subject position. “Which band played their first gig at the Boardwalk club?” → Oasis.
The research found: when a query reverses the subject/object order, recall gets noticeably harder.
In open-ended generation (recall), reverse questions are consistently harder than direct ones. But in multiple-choice verification (recognition), reverse questions are no harder — often easier.
That dissociation is the key insight: if a model can recognize the right answer among distractors but can’t generate it when the query direction flips, the fact is encoded — even recognizable — but hard to recall because the question departs from how the fact was encountered during training.
In other words: the reversal curse is a recall problem, not a missing-knowledge problem.
Two more findings worth your attention:
1. Rephrasing doesn’t help. Reordering does. The researchers tested whether rewording questions changed recall. It didn’t, not significantly. What mattered was whether the subject/object order matched.
2. Long-tail facts aren’t “unlearned” — they’re unreachable. Comparing low-popularity and high-popularity facts, the encoding gap was modest. The recall gap was much larger. The long-tail problem gets reframed: rare facts aren’t absent from the model’s parameters. They’re present, but difficult to access.
How much does thinking mode recover?
The research also tested “thinking” — letting the model do intermediate computation before the final answer (chain-of-thought prompting, thinking-optimized models).
The results:
- For encoded-but-not-directly-recallable facts, thinking recovers 40-65%
- For never-encoded facts, thinking recovers only 5-15%
That gap is exactly what you’d expect if thinking acts primarily as a recall-facilitation mechanism — it helps the model reach knowledge it already has, rather than deriving answers through complex reasoning on the spot.
Thinking also narrows both gaps: between popular and rare facts, and between direct and reverse questions.
But it’s not free: it’s computationally expensive, and nobody knows when to trigger it — not the models, not the researchers. And scaling is explicitly ruled out as the fix: in the Gemma 3 family, bigger models show far fewer encoding failures, but recall failures persist and grow as a share of remaining errors.
Scaling improves storage. It doesn’t improve access.
What this means for independent site SEO
Now for the part you actually care about. In SEJ’s analysis, Roger Montti put forward an idea I fully agree with:
Intuitively, it may be beneficial to order subject and object entities according to the most common way that queries order them.
Note: that’s a reasonable hypothesis, not a paper finding — the research never claims entity ordering will make an LLM pick your page. But the logical chain for SEO is compelling:
AI answers are recalled from your content → recall is shaped by the subject/object order encountered in training → the order you write in is the order AI trains on → writing facts in the order users ask them makes it easier for AI to “recall” your content.
Concretely, here’s what I’d do:
1. Cover both directions of every key fact
Don’t write a critical fact in only one direction.
Say you’re writing “Shopify supports the TikTok sales channel.” Beyond that subject-first sentence, cover the direction users actually ask: “Which platforms support TikTok sales?” The ideal approach is to write both naturally in the same article:
- Shopify integrates the TikTok sales channel
- Want to sell on TikTok? Use a platform like Shopify that supports the TikTok channel
That’s the direct-question + reverse-question double coverage, applied to content.
2. Pin entity relationships in the headline and first paragraph
AI recall and citation lean hardest on titles, opening paragraphs, and the first sentence under a heading. Putting “what A is and how A relates to B” in the most prominent spots gives the model’s retrieval the easiest handle.
3. Organize content around user phrasing, not encyclopedia phrasing
Users typically put the answer in the subject position (“who…”, “which…”, “what brand…”), while encyclopedia-style writing puts it in the object position. Write with the user’s most common question direction as your primary expression — don’t just follow the reference-book word order.
4. Write long-tail facts extra clearly
The research says rare facts are stored but hard to access. The translation: the niche facts you specialize in are exactly the differentiated content AI is most likely to pull from you — provided they’re written clearly. Use the simplest “A is B” structure and nail the entity relationship in the first paragraph. No meandering subordinate clauses.
Neo’s take
1. This is the first paper-level evidence that phrasing order affects AI answers
For two years we’ve talked about AEO, GEO, and content structuring — mostly from experience. This paper supplies the mechanism: the order of entities in training text genuinely shapes later recall. For anyone who writes content, “how you write” finally has a scientific basis.
2. It turns entity SEO from mysticism into mechanics
Entity SEO has always felt vague — “just make sure Google knows you, your products, and your relationships.” This research gives it a concrete physical mechanism: subject/object order. From now on, every product description, FAQ, and comparison post gets a checklist question: does the subject/object order of this sentence match how users ask the question?
3. Don’t over-read it, but act now
Three cold showers:
- The research is based on WikiProfile (2,150 Wikipedia facts); Wikipedia’s prose style is very different from commercial site content
- “Order entities the common way” is the researchers’ intuitive suggestion — the paper never validates a direct effect on indexing or citation
- This is about model recall; how models pick citation sources from the web is a separate mechanism
But the case for acting now stands: bidirectional coverage, user-phrasing-first, pinning entity relationships in the first paragraph — these cost almost nothing, and even if the research is later overturned, these practices can’t hurt your normal SEO. It’s the classic “limited downside, unlimited upside” play.
4. A bigger signal: the direction of AI-friendly writing
This paper hints at something larger: if subject/object order affects recall, then “AI-friendly content” and “content that reads well to humans” are starting to diverge in subtle ways. I expect writing tools that optimize for model recall — automatically ordering entities to match how people ask about your target keywords. The people who think through this logic first will eat first.
Wrapping up
“Empty shelves or lost keys?” Google’s answer: the shelves are stocked. The keys are the problem.
For independent site owners, that’s both a warning and an opportunity:
- Warning: in the AI era, the competition isn’t about who knows more — it’s about whose knowledge AI can actually retrieve
- Opportunity: the order, structure, and placement of entities in your content finally has a verifiable optimization direction
Starting today, before you write any critical fact, ask yourself one question: how would a user ask this? Did I write it the way they’d ask?
I’m Neo, see you in the next one.