Same Number, Seven Answers: Can You Actually Trust the AI Crawl-to-Refer Ratio?


Hi everyone, this is Neo.

Here’s a mildly dark fact to start with: over the past 13 months, one metric from one company has been quoted in seven different versions, and the highest and lowest of them differ by roughly a factor of 30. Every single version ended up in a slide deck, a contract negotiation, or a “should we block this AI crawler” decision.

The metric is the crawl-to-refer ratio: how many pages an AI platform takes from your site versus how many visitors it sends back. For Anthropic, the published figures include 70,900:1, 38,000:1, 23,951:1, 11,122:1, 10,300:1, 4,580:1 and 2,237:1 — all attributed to Cloudflare. Two of them claim to cover the same month, and they differ by a factor of 17.

The interesting part isn’t that someone lied. Quite the opposite: Cloudflare published the formula in full and, unusually, disclosed the limitations of its own method. The problem is what happens between the publisher of a number and the person deciding on it — things get dropped. So let’s open this metric up: how it’s built, why it scattered across an order of magnitude, and how an independent site should use it rather than be used by it.

1. The formula, plainly

Cloudflare launched the metric in July 2025 and now shows it on Radar. The math is short enough to memorize:

Numerator: total requests from user agents associated with a given platform where the response was Content-type: text/html. Denominator: total HTML requests whose Referer header contained a hostname associated with that same platform. Normalize to a single referral.

In plain English: it pulled 70,000 pages from you and sent back one visitor — that’s 70,900:1.

In the original launch table, for the week of June 19–26, 2025, Anthropic came out at 70,900:1 while Mistral came out at 0.1:1 — meaning Mistral sent back ten visitors for every page it fetched. Same table, two extremes, a gap of 700,000x.

The metric went viral because it hit a nerve everyone shares: legacy search crawlers take your pages and send visitors back, and that trade is what funds publishing on the open web. AI systems answer in place, so the taking continued and the sending back largely didn’t. A ratio you can put in a report is far more useful in a board meeting than a general feeling that you’re being ripped off.

2. Four denominators wearing one coat

This is the core of the piece. A ratio is a numerator over a denominator, and this denominator has four different things stacked inside it that almost nobody carries forward when quoting the figure.

The window

The launch post used June 19–26, 2025 — one week. In a separate post the same month, Anthropic was 73,000:1, OpenAI 1,700:1, and Google roughly 14:1. Same publisher, same month, different window, different number.

Cloudflare itself reported Google’s ratio moving 19.4% week over week, purely because Googlebot crawling dropped from June 24 onward. Read that again: one crawl scheduling decision moved the published number by a fifth inside seven days.

Today you can find quarterly figures, monthly figures, rolling 28-day figures and single-week figures circulating side by side, looking like the same measurement.

Which bot got counted

Cloudflare states that a platform’s training crawler and its user-request crawler run under different user agents but are aggregated under a single platform name in the analysis.

Those two behaviors have nothing in common. One consumes at scale and returns nothing by design; the other fetches on demand and can produce a citation. Rolled together, the number describes neither. Worse: some operators run purpose-split crawler fleets and others run unified ones, so comparing a platform figure against Google is structurally unsound — you’re measuring a different kind of object on each side.

Cloudflare’s own by-purpose data shows how unstable that bucket is: for July 1–28, 2025, about 80% of AI bot crawling was training traffic. By June 2026 training was down to 52%, while a mixed-use category held more than 36%.

The denominator isn’t referrals — it’s referrals that announced themselves

This one is the most damaging, and Cloudflare wrote it down in the launch post itself: traffic referred by Claude’s native app (and, they believe, other providers’ native apps) arrives without a Referer header, and the denominator only counts requests that carry one. So the published ratios “may overstate” the imbalance — “but it is unclear by how much.”

Unclear. Their words.

One third-party analysis puts it at up to 70.6% of AI traffic arriving with no referrer data at all, and — this is the twist — that invisible slice reportedly converts at around 10.21%, versus 1.66% for visible AI traffic and 0.15% for organic search. So the most valuable visitors are precisely the ones this metric cannot see.

And the error only runs one way: every unrecorded referral inflates the ratio, and nothing deflates it. A metric with one-directional measurement error is fine for spotting extremes and unfit for close calls.

The collection boundary

Cloudflare deliberately excludes referral traffic from Google’s own network, on the grounds that Chrome’s speculation-rules prefetching isn’t a person consuming content. That’s defensible — but it’s a judgment call. Another analyst could handle prefetch traffic differently and produce a different Google figure, and neither analyst would be wrong.

3. Why this matters: the number is already making irreversible calls

If this were just an industry nitpick about rigor, it wouldn’t be worth a post. The problem is what the ratio is being used to decide:

  • Publishers use it to choose which AI crawlers to allow and which to block.
  • Marketing teams use it to argue that AI referral traffic isn’t worth pursuing.
  • And it has been productized: it’s in Cloudflare Radar, and in August 2026 Microsoft added an “AI Scrape-to-Referral Ratio” card to Clarity’s Bot Analytics dashboard. When a number sits in a dashboard, right next to a bot’s name, teams read it as a verdict on that bot. It isn’t built to carry that weight.

Two consequences you should know about:

One: the ranking flips. In the rolling 28-day window ending July 21, 2026, Mistral “beat” Anthropic at 3,389:1 versus 2,237:1. Slice Cloudflare Radar a different way and Anthropic is on top again. Same network, different cut, different “most extractive platform.”

Two: it measures the venue, not the contest. As one analysis put it nicely: on the same engine in the same window, two brands can differ 4x in citations and still report near-identical crawl-to-refer ratios — because the number is mostly a property of the engine’s architecture and caching, not of your content. So can you benchmark against competitors with it? No. Nobody publishes a cross-site benchmark for this metric, because there isn’t one to publish.

The inverse holds too: Anthropic’s ratio falling from 70,900:1 to under 2,000:1 in 2026 wasn’t mainly sites getting better at earning referrals — it was operators adding caching and shipping consumer products that actually link out. The number moved, and it had nothing to do with you.

4. Three places where the metric genuinely earns its keep

I’m not saying it’s worthless. Used correctly, it’s good:

Use case How it helps Watch out
Cost accounting Server load, bandwidth and cache-miss traffic are real money. One publisher reported saving roughly €42,000 a month in CDN bandwidth after blocking unauthorized crawlers This belongs to finance and infra, not to a marketing KPI
Licensing leverage In content-licensing talks, crawl volume is the one number both sides accept It’s a pricing number, not a management number
Anomaly detection A crawler hammering your whole catalog overnight shows up instantly in your own logs Use your logs, not a network-wide average

Keep one sentence in mind: the ratio answers “how much is this bot taking,” not “is this bot worth allowing.” Only the second question has a business answer.

5. The operator’s checklist: ask three separate questions

“Should I block this crawler” actually hides three different questions, and one ratio answers none of them:

1. Cost is a server question. Pull crawl volume from your own access logs and act on measured load — rate limiting is proportionate and reversible. Don’t let a network average make the call for you.

2. Visibility is a citation question. Run your own prompt set against the assistants and measure citation share, not referrals. For an assistant that answers in place, near-zero referrals is the expected steady state, not a failure. A formula that scores an uncited crawl and a well-cited answer with no click as the same loss cannot tell an unknown brand apart from the dominant answer source.

3. Access is a policy question. Use the purpose layer in robots.txt so training and retrieval can get different answers. That’s the only path that keeps the discoverability half of the trade while declining the other. One lock for both purposes usually costs you both.

Here’s a free filter to use from now on. Whenever someone hands you a measurement, ask three things: what period does it cover, what got grouped to produce it, and where did collection stop? If those answers aren’t immediately available, they haven’t given you a number — they’ve given you a shape that resembles one.

Neo’s take

Every time a new intermediary inserts itself between you and your audience, the first metric to appear is the one the intermediary can compute most cheaply — and that metric describes the intermediary’s business, not yours.

Run the history: server logs, referrer headers, impressions, average position. Each arrived as somebody else’s accounting and got adopted as our scorecard. The crawl-to-refer ratio is just the newest entry in that lineage — and it won’t be the last.

For independent sites, three rules I run by:

  • Don’t use “there’s no trustworthy industry ratio” as an excuse to do nothing. Your own logs are the best data source you have, and bandwidth plus server load are things you can genuinely optimize.
  • And don’t use “this ratio is huge” as a reason to block everything. You might be shutting the exact retrieval path that would have cited you. Crawling is a cost; absence is a bill — one that never arrives and can’t be reconciled.
  • The only metric worth watching is a set: how often AI cites you, mentions you, and uses you as a source in the questions that matter to your business. It’s an uglier number, but it’s honest — and it’s the only one that moves because you made your content better.

One last line for the content people: this industry is busy packaging number-anxiety and selling it back to you. When someone hands you a figure, you should be able to name its window, its grouping and its boundary. If you can’t, you didn’t buy insight — you bought something shaped like a number.