ChatGPT's Search Index Isn't Abandoning Small Sites: What the New Data Really Means
Hi everyone, this is Neo.
Lately there’s been a wave of anxiety sweeping through the independent site community: “AI search only recognizes big brands — small sites are finished.” It got worse last month when a well-known technical SEO guy dug into ChatGPT’s network traffic and found a field called labrador in the results — packed with names like Reuters, WSJ and Wikipedia. The conclusion everyone jumped to: ChatGPT’s in-house index is basically an allowlist, and small sites can’t get in.
Well, I’ve got some good news for you today. New research from Resoneo, a French SEO consultancy, directly challenges that narrative: ChatGPT’s in-house search index serves small websites too — and it serves them exactly the same way it serves big publishers with content licensing deals.
In this post, I’ll walk you through the raw data, the mechanics behind it, and what this actually means for independent site owners. No fluff.
Rewind: Where Did the Panic Come From?
It started in July. SEJ contributor Suganthan Mohanadasan published a viral piece called “How ChatGPT Actually Picks Sources,” where he intercepted ChatGPT’s browser network traffic (the JSON data) and uncovered the secret behind its citations.
Every search result carries a result_source field with one of four values:
| Field value | What it is | Where it shows up |
|---|---|---|
labrador |
OpenAI’s own index | News, reference content |
serp |
Baseline open-web fetching | News (Yahoo, StreetInsider, etc.) |
bright |
Bright Data (commercial scraper) | Shopping, finance, weather, local |
oxylabs |
Oxylabs (rival scraper) | Regional and local press |
His initial read: labrador is an allowlist of established publishers like Reuters, The Guardian, WSJ, FT and Wikipedia — outlets that have signed content deals with OpenAI. His words: “it isn’t one you get into unless you own a national newspaper.”
Cue the meltdown. Independent site owners were already worried about AI search squeezing their traffic — now they couldn’t even get cited?
But credit where it’s due: Suganthan was careful to flag this as a small sample from a single account — directional, not measured. And he published a correction in July after digging deeper.
What Resoneo’s Research Actually Found
Resoneo took a harder-core approach: they used their own Chrome extension to capture 1,249 ChatGPT answers in July, reading directly from ChatGPT’s server data stream, which tags every web result with the name of the pipeline that fetched it.
Three findings matter:
1. Licensing deals change nothing about how a page is served
When Resoneo compared pages served through the labrador pipeline, they found: whether or not a site has a content deal with OpenAI, it’s served identically — same format, same snippet length, same freshness.
In other words, labrador isn’t a paid allowlist. It’s OpenAI’s own search index, topped up with press feeds and open science archives. OpenAI accesses it directly without paying a third party — that’s what sets it apart from the commercial scraper pipelines.
2. In free accounts, the in-house index is the workhorse
In Resoneo’s free-account data:
- Questions with settled answers, local businesses, and product queries came through
labradoralmost every time - News results split roughly evenly between the in-house index and Google scraping
For everyday free users, ChatGPT’s own index is the main pipeline.
3. In paid thinking mode, Google scraping dominates
On paid accounts in thinking mode, Resoneo recorded 16,407 search results, and:
- About 75% came from Google scraping
- About 24% came from the in-house index
That’s the fascinating part: ChatGPT dynamically mixes pipelines depending on account type and usage mode. In deeper paid thinking mode, it leans even harder on Google’s open index.
What This Means for Independent Sites
Let’s connect the dots — in three layers, because there’s more here than meets the eye:
Layer 1: Small sites genuinely can get into ChatGPT’s index
This is the headline, and it’s a relief. Resoneo was explicit: “hundreds of outlets with no OpenAI content deal were served by OpenAI’s in-house search index exactly the way its licensed partners were.”
The “ChatGPT index = big media allowlist” narrative? Officially dead. labrador covers the open web, not a closed licensing pool.
Layer 2: But “getting in” and “getting picked” are different things
The data also reveals a sobering reality: in paid thinking mode, 75% of ChatGPT’s results come from Google scraping.
What does that tell us? In many scenarios, ChatGPT is essentially borrowing Google’s search results. If you want to show up in AI search, ranking in Google first is still the fundamental skill — which lines up with what we’ve been saying all along: Google SEO is not dead in the AI era.
Layer 3: Different pipelines demand different strategies
labrador (in-house) and bright/oxylabs (commercial scrapers) cover different content ecosystems:
- In-house index: rewards structured, parseable, factually solid content — press-style releases, reference material, product information all have a shot
- Commercial scrapers: pull open web pages — Reddit, industry reviews, e-commerce detail pages are all in there
So don’t put all your content eggs in one pipeline. Structured data, clear entity information, and machine-parseable page architecture are the underlying capabilities that work across every pipeline.
What to Do Now: Four Actionable Steps
Enough analysis — here’s what you can actually do:
1. Stop worrying about “signing a deal with OpenAI”
Some owners panicked at the idea that you need a content licensing agreement to get into ChatGPT, and scammers are already selling “channels” to do it. This data says plainly: no deal needed — you’re served the same either way. Spend your energy on content quality instead.
2. Get “indexed by Google” right, first
75% of results come from Google scraping — that sentence deserves to be on your wall. If Google can’t find your page, AI search won’t either. Audit the fundamentals: robots.txt isn’t accidentally blocking crawlers, sitemaps are submitted, internal links are sane, page speed is decent.
3. Verify your own AI visibility with tools
Resoneo offers its Chrome extension for free, and Suganthan’s article includes a browser console script that reads the result_source field. You can literally see which pipeline fetches your own pages in ChatGPT’s results. Data over vibes.
4. Make your content machine-readable
Since labrador is an in-house index, how OpenAI reads your page matters: clear heading hierarchy, answer-first opening paragraphs, structured product specs, FAQ blocks. All of these raise your odds of being indexed and cited by AI.
Neo’s Take
The truly valuable thing about this research isn’t the headline “small sites can get into ChatGPT’s index” — satisfying as that is. It’s the real operating mechanism it reveals about AI search.
Here’s the thing: most of the industry’s understanding of AI search comes from black-box testing. Researchers fire thousands of prompts, tally which brands appear in answers, and produce share-of-voice reports. That approach only ever sees outputs, never mechanics.
Resoneo and Suganthan went the other way — reading the server data stream directly to see who fetched each result and through which pipeline. It’s like going to a restaurant where everyone else just tracks which dishes are most popular, while these guys sneak into the kitchen to see which stove cooked them.
My honest read:
First, the “allowlist anxiety” can be retired — but “mediocre content anxiety” should be adopted. The index is open, and so is the competition. Before, you were competing with big publishers for Google rankings. Now it’s about “whose pages are easier for AI to understand and trust.” That’s actually friendlier territory for small sites — because the big-brand halo counts for less than you’d think in AI’s eyes.
Second, multi-pipeline strategy becomes the new normal for independent sites. ChatGPT uses different pipelines in different scenarios, and Perplexity and Gemini will likely follow. Don’t bet everything on “getting cited by AI.” Google rankings, direct traffic, email lists, brand search — work every channel.
Third, and most important: data transparency is killing “vibe-based SEO.” SEO used to be a black art where anyone could claim anything. Now people can read ChatGPT’s pipeline allocation directly, and more reverse-engineering research like this is coming. Whoever understands the mechanics first gets the advantage. It’s time to make “understanding the mechanism” a habit.
Stop panicking, start working. If your content is genuinely good and your structure is clear, ChatGPT’s index will find you eventually.