How Much Traffic Did AI Overviews Actually Take From Wikipedia? The UW Paper Cut Its Own Estimate From 15% to 5% — Google Still Disputes It, But the Real Story Is Elsewhere


Hi everyone, this is Neo.

Today I want to talk about a research paper — but the interesting part isn’t the number. It’s the fact that the number got revised.

Mehrzad Khosravi and Hema Yoganarasimhan at the University of Washington published a working paper called “Impact of AI Search Summaries on Website Traffic.” They used Wikipedia as the test case to estimate how much traffic Google’s AI Overviews (AIO) actually take away from publishers.

When the paper first appeared in February, the headline finding was a 15% drop. By the latest version, updated September 2, that had become a 5.45% / 4.82% drop.

Same authors. Same dataset. Same causal design. The estimate shrank to about a third of its original size. And yet most of what’s still being quoted in the industry today is the original 15%.

Before you roll your eyes at academia revising itself again — this revision is worth reading closely, because it exposes the weakest link in the whole “AI is stealing traffic” conversation: measurement.

1. Why the study design is genuinely clever

The hard part of measuring AI’s traffic impact is that you can’t run a controlled experiment. You can’t show AI Overviews to half your users and hide them from the other half.

So the smart move is to find something that already comes with a control group. Wikipedia happens to fit perfectly:

  • The same article exists in dozens of language editions, covering essentially the same content;
  • Different language editions are read mostly by people in different regions;
  • Google turned AIO on by default in the United States in May 2024, while German and French readers didn’t have it by default during the sample period.

That gives you a natural quasi-experiment: the same article, treated differently. If English referrals fell relative to German after May 2024, the gap is your estimate.

The data is Wikimedia’s public article-level clickstream (monthly, counting human referrals from external search engines), covering December 2023 through December 2024. The matched samples include 499,927 English–German article pairs and 530,873 English–French pairs, estimated with a Poisson pseudo-maximum-likelihood difference-in-differences model built for count data.

The results:

Comparison Change in English search referrals Scale
English vs. German -5.45% about 100.27 million fewer search-originated visits per month
English vs. French -4.82% roughly 1.2 billion fewer per year
English vs. Japanese (supplementary) -16.53% shorter window, different control group; authors call it directional support only

2. How the 15% appeared — and where it went

This is the part I find most interesting.

Version 1: used daily pageviews, with Hindi, Indonesian, Japanese and Portuguese as controls. Result: -15%, plus a hair-raising extrapolation — the English articles in that sample were losing “11.5 million visits per day, or about 4.21 billion per year.”

Version 6 (September 2): switched to monthly search referrals, with German and French as controls (a stronger choice, since those readerships sit mostly outside the U.S. and had no default AIO during the sample). Result: -5.45% / -4.82%.

Three changes stacked up, and the number shrank:

  1. The metric changed. Pageviews include internal navigation, direct visits, social and bookmarks. Search referrals only count what search engines send. AIO is a search feature, so search referrals are the cleaner causal channel — but the switch also squeezes out the noise contributed by every other entry point.
  2. The control group changed. Using Hindi and Indonesian as controls drags in whatever else was happening to the English-language internet in 2024. German and French are better comparators, and the fact that both give consistent answers (-5.45% and -4.82%) actually increases confidence.
  3. Daily granularity became monthly. The finer the slice, the more AIO’s appearance and disappearance rattles the estimate.

I wrote something in my piece on the crawl-to-refer ratio that applies verbatim here: change the window, change the definition, and the same phenomenon produces numbers that differ by multiples — with nobody technically being wrong.

3. Why Google pushes back, and what the paper admits

Google’s objection is specific, and it isn’t bad faith: Wikimedia’s public clickstream data lumps all external search engines together. So the analysis can’t isolate Google. That’s what Google told UOL Tilt.

The authors’ rebuttal holds up too: Google is roughly 90% of the global search market, so the other engines are a rounding error; and the design uses the May 2024 default-rollout as its timing break, not a shift in Google’s market share.

But what I appreciate most is the paper’s own list of limitations. It’s honest to the point of being reassuring:

  • It counts human referrals directly from search engines only. If someone arrives from Google and reads three more articles, that’s one visit;
  • Most English Wikipedia readers are outside the U.S., so the measured effect is diluted;
  • The design can’t fully rule out other 2024 events that hit English search but not German or French;
  • The “a comparable ad-supported site would lose $10.82M to $37.08M a year” figure is purely hypothetical — Wikipedia runs no ads, so that money doesn’t represent anything real moving to Google.

That last one matters. “Losses of $X billion” always travels further than “an estimated confidence interval.” Next time you see it, ask who the money belonged to and how it was calculated.

4. The story only makes sense with two other datasets next to it

Reading the paper alone will mislead you. Two other threads have to be read alongside it.

Thread one: Wikimedia’s own numbers (October 2025, Marshall Miller, Senior Director of Product).

  • From March to August 2025, human pageviews across Wikipedia fell about 8% year over year — a figure confirmed only after improved bot detection was used to reclassify a wave of disguised traffic from Brazil;
  • Bad bots accounted for 37% of total traffic (up from 32% in 2023);
  • Bots made up 65% of the most expensive traffic hitting the core data centers, and about 35% of pageviews;
  • The team blocks or throttles roughly 1.5 billion requests per day;
  • Bandwidth for downloading media files from Commons is up 50% since January 2024.

Note Miller’s language: he attributes the decline to generative AI and social platforms, but explicitly frames that as a belief, not a causal finding.

Thread two: Pew’s click behavior study.

Pew analyzed 68,879 Google searches from March 2025, captured from the browsing activity of 900 U.S. adults:

  • When a search included an AI summary, people clicked on traditional results 8% of the time, versus 15% without one;
  • Only 1% of visits to a page with an AI summary ended in a click on a source inside that summary;
  • Wikipedia, YouTube and Reddit together accounted for 15% of cited sources.

Thread three: Google’s own position.

Liz Reid, Google’s head of Search, said in an August 2025 blog post that total organic click volume stayed “relatively stable” year over year — without publishing a number — and repeated that framing on Bloomberg’s Odd Lots podcast in April.

Stack them together:

Human referrals are falling (Wikimedia’s official -8%; the paper’s estimated -5% for search referrals), machine demand is rising (bandwidth +50%, 65% of the most expensive traffic is bots, 1.5 billion requests blocked per day), and the platform says “total clicks are stable.”

All three can be true at once. Because they’re measuring three different books.

5. Averages lie: -19.6% and -7.4% are the real headline

If the paper only gave you a single -5%, it would be easy to shrug off. “Five percent on average” doesn’t feel like a crisis.

But the paper breaks it down by content type. From the earlier daily-data version:

Content type Relative decline
Culture -19.6%
Geography -16.6%
History roughly -14%
STEM -7.4%

In plain English:

Informational content that a short answer can fully satisfy gets hit hardest. Content that requires multiple steps, specific parameters, or a procedure holds up.

Wikipedia’s Culture entries are the perfect example of “you ask, it answers” — an AI Overview can resolve the query in place. STEM entries tend to involve formulas, derivations and specific conditions, so people still go read the page.

This matches every practitioner’s gut feeling about what AI can replace — except now there’s causal evidence behind it.

6. What independent site owners should do with this

1. Run a “Culture-article audit” on your own site.

Pull your informational pages and ask one question of each: could AI answer this page’s core point in three sentences? If yes, that page lives in the danger zone. You don’t have to delete it — just stop betting your KPIs on its clicks.

2. Keep two separate books.

  • The human referral book: organic clicks in Search Console, impressions in the AI reports;
  • The machine demand book: AI crawler hits in your server logs, bandwidth cost.

Watch only the first and the world looks like it’s collapsing. Watch only the second and everything looks fine. You need both to see reality.

3. Design an exit for “cited but not visited.”

Wikipedia’s dilemma is a template: cited the most, referred to the least. Its answer isn’t an anti-AI technical patch — it’s turning citations into a different kind of entry point: direct brand searches, donations, Wikimedia Enterprise, community.

Same for your site. If your pages show up in AI answers but get few clicks, the question isn’t “how do I make them click.” The question is: someone now knows my brand name — where do they go next, and am I waiting there?

4. Don’t use this paper to convince anyone that “AI stole X%.”

It will change. This is its sixth revision, and most of the “landmark studies” the industry keeps quoting stopped at version one.

Neo’s take

When I saw the 15% → 5% revision, my first thought wasn’t “academia is unreliable.” It was — if this paper had never been revised, the 15% would have followed us around for years.

How many of the “hard numbers” our industry cites endlessly are frozen at some author’s first draft or some PR team’s first estimate? We watched the exact same show in my piece on the crawl-to-refer ratio: one metric, seven published versions, a 30x spread, and every single one of them ended up in a deck.

So the real takeaway here isn’t the 5%. It’s two things:

First, AI is taking traffic, but it’s a diversion, not an extinction. For Wikipedia — a site cited constantly by AI, with the most standardized content on the web — a full year produced roughly 5% fewer search referrals and 8% fewer human pageviews. Not halved. Not zeroed. Most “traffic down 60%” stories happen to sites that already had problems. Blaming AI is far more comfortable than admitting your content was thin.

Second, the 12-point gap inside the average is your opportunity. The distance between -19.6% and -7.4% isn’t luck. It’s content type. Whatever a short answer can resolve is going to get eaten. Whatever needs experience, parameters, steps and judgment is still standing.

My rule hasn’t changed: content that can be fully answered in three sentences shouldn’t be expected to bring clicks. Content that can’t be is my battlefield. Move your resources toward the second group and be patient — because the stronger AI gets, the shorter the list of things that “can’t be answered in three sentences” becomes. And whatever survives on that list will be worth more than ever.