Bing's AI Citation Tracker, the Hidden HTTP 'Ghost' Homepage, and the Truth About Google's Crawl Limits


Hi everyone, this is Neo.

In independent-site SEO, the scariest thing is groping the elephant in the dark — guessing at things we can’t see.

We never knew whether AI was actually using our content — just blind guessing. Sometimes the site name in search results mysteriously changes, and you dig through all your code without finding the cause. Lately, word spread that Googlebot only crawls the first 2MB of a page, and everyone panicked that their long-form articles were a waste of effort.

This week, those black boxes finally opened. Bing launched an AI citation dashboard, Google’s John Mueller exposed the ghost-homepage trap, and fresh data showed us: relax, 2MB is actually a lot.

Today I’m doing a deep recap of the three big SEO stories from this week — and what they actually mean for independent sites.


1. Bing Webmaster Tools: Finally, You Can See How AI “Cites” You

Easily the biggest news of the week. Microsoft officially launched the AI Performance Dashboard in Bing Webmaster Tools (BWT).

1. What Can You See?

Before, Search Console only showed clicks and impressions. Now Bing gives us a glimpse beneath the surface of Generative Engine Optimization (GEO). The dashboard shows:

  • Citations: how many times your site gets cited in answers generated by Copilot and Bing AI.
  • Grounding queries: this is the big one! It shows which user search queries AI used to “locate” your content when generating an answer.

2. Neo’s Take

Google lumps AI Overviews data into Search Console without breaking it out separately. Bing just leapfrogged them.

The dashboard doesn’t show click data yet — so you can’t tell how much traffic those citations bring. But grounding queries alone are enormously valuable.

In practice: Head into BWT and check what terms AI uses to find you. Say you sell “Outdoor Furniture.” You might discover AI cited your review article via the long-tail query “best weather-resistant patio set.” That hands you an optimization direction: answer those specific questions more explicitly in your content to boost the odds AI “picks” you as its answer source.


2. Why Is Your Site Name All Wrong? Beware the Hidden HTTP “Ghost” Homepage

Recently, Google’s John Mueller shared a truly spooky case: It’s like you’ve put up a brand-new sign, but Google insists on showing your old, battered one.

1. The Crime Scene

A site had fully migrated to HTTPS, yet its site name and favicon showed up wrong in Google’s search results. The owner checked every page — browsers loaded HTTPS fine, and the structured data in the code was correct.

2. Who Dunnit?

Turns out, the server still had an HTTP version of the homepage (note: not HTTPS). When you visit http://yourdomain.com in Chrome, the browser “helpfully” auto-redirects you to https://yourdomain.com or upgrades the request. So as a human user, you never see that old HTTP page.

But! Googlebot doesn’t auto-upgrade like Chrome. If your server allows HTTP access without a 301 redirect, Googlebot can crawl that stale, misconfigured HTTP page — and extract the wrong information from it.

3. Neo’s Practical Advice

Don’t trust what your browser shows you. Browsers hide a lot of technical flaws in the name of user experience. When you hit one of these “ghost” problems, go straight to the command line with curl.

Run this in your terminal:

curl -I http://yourdomain.com
  • Correct result: 301 Moved Permanently, pointing to the HTTPS version.
  • Wrong result: if you get 200 OK, congratulations — you’ve found the ghost page. Go add a forced HTTPS redirect in your server config right now.

3. The Googlebot 2MB Crawl Limit: Should You Really Worry?

Google recently updated its docs to state that Googlebot only reads the first 2MB of HTML and resource files when crawling (64MB for PDFs). That set the SEO world on fire: “If my page exceeds 2MB, is everything after it worthless?”

1. Let the Data Speak

SEO expert Roger Montti analyzed real data from HTTP Archive, and the results are reassuring:

  • The median mobile page HTML size is just 33KB.
  • Even the top 10% “heaviest” pages come in at only 151KB.
  • That’s nowhere near the 2MB (2048KB) limit.

2. Who Actually Gets Hit?

For 99.9% of independent sites, this isn’t a problem at all. Only pages with genuinely terrible code could exceed the limit. For example:

  • Inline images: base64-encoding multi-megabyte high-res photos straight into the HTML.
  • Giant JSON blobs: stuffing massive JSON data into the HTML for front-end rendering.

3. Neo’s Summary

Don’t panic — but stay disciplined. 2MB is plenty, but that doesn’t mean you can waste it. Keeping your code clean has always been SEO 101. If your HTML really does exceed 2MB, SEO is the least of your worries — can your users even load that page? It’d be slow as molasses!


Summary

This week’s SEO signals are clear: transparency is rising and tools are getting sharper.

  1. Sign up for Bing Webmaster Tools — even if you only care about Google, Bing’s AI data is a goldmine of content ideas.
  2. Check your HTTP redirects — don’t let a ghost page wreck your brand image.
  3. Keep your code lean — the 2MB limit only punishes bad code; good sites have nothing to fear.

SEO is a marathon. Stay sensitive to technical details, and you’ll stay standing in the AI era.