A robots.txt rewrite to stop AI bots took a site out of Google

Post from r/TechSEO: Blocked Googlebot variants by mistake. Need
Post: r/TechSEO, u/khawarmehfooz

Summary

Google crawls with several user-agent tokens, and a robots.txt group that names one of them overrides the general Googlebot group. An allowlist written in a hurry can leave the HTML crawlable while cutting off Google Images and Search Console's live URL test, and the site owner will not see the problem until pages vanish.

Audit robots.txt by variant rather than by searching for the Googlebot string, and delete variant groups you did not mean to write. Handle a request flood at the server or WAF, not by rewriting the whole file.

A site owner who blocked most crawlers in robots.txt after an AI company’s bot sent more than 50,000 requests in 24 hours says the site then dropped out of Google entirely, including searches for its exact brand name. The account is in a thread on r/TechSEO from u/khawarmehfooz, who runs an image-heavy marketplace site for a mobile game in around 20 locales. The rewritten file allowed Googlebot, and the poster lists Googlebot-Image and Google-InspectionTool among the Google crawlers it blocked by accident.

Google crawls with several user-agent tokens, not one. Google’s list of its common crawlers documents Googlebot-Image, Googlebot-Video, Googlebot-News, Google-InspectionTool and GoogleOther as tokens that can be addressed separately, each in its own robots.txt group. A file that allows Googlebot and disallows Googlebot-Image keeps the HTML crawlable while cutting off Google Images, which on a site whose pages are mostly images is much of what Google has to rank.

# Pages stay crawlable, Google Images does not
User-agent: Googlebot
Allow: /

User-agent: Googlebot-Image
Disallow: /

The thread’s account of how that happened does not quite match Google’s documentation. Google lists Googlebot as an alternative token for both Googlebot-Image and Google-InspectionTool, so an allow group for Googlebot plus a catch-all Disallow would still have let both through. For them to have been blocked, the file most likely named them. The risk in a copied block-bad-bots list is not that it omits Google’s variants, it is that it names them. Two variants do not fall back to the Googlebot group, so a catch-all rule does stop those.

TokenAlso obeys Googlebot rules
Googlebot-ImageYes
Googlebot-NewsYes
Google-InspectionToolYes
Googlebot-VideoNo
GoogleOtherNo

Blocking Google-InspectionTool costs something different. That token is what URL Inspection in Search Console uses for live tests, so disallowing it removes the one check that would have shown the problem while it was happening.

Recovery a month after the fix is partial. Google began indexing the site again in early August, per the thread, and clicks peaked in mid-August at roughly 20 a day. Traffic then crashed around 24 August and has stayed close to zero since. Search Console shows 167 clicks in total against 2.18k indexed pages.

The poster asks whether the old robots.txt block could still be suppressing rankings. The likelier answer is no. Robots.txt is not a penalty, and once the file allows a crawler it stops being the constraint on that crawler. A more likely explanation for the August peak and the crash sits in the pages themselves. The thread’s only reply, from u/cinemafunk, says the site carries almost no text explaining what it is or what it is for, and pages that are mostly images give Google little to rank once a fresh reindex settles. That points the work somewhere other than the robots file.

The block that replaced the emergency one now targets AI-training and SEO-tool crawlers, and that category is looser than it sounds. A list built to stop training crawlers such as GPTBot and ClaudeBot often carries retrieval bots like OAI-SearchBot and PerplexityBot too, and those are what fetch pages for citations in AI answers. Blocking SEO-tool crawlers also cuts off the site audits and backlink data from tools you may pay for. Broad AI rules reaching Google is not hypothetical either: Cloudflare’s managed AI-training block has caught Googlebot as a side effect.

What to do

  • Open Settings, then robots.txt in Search Console to see when Google last fetched the file, whether the fetch succeeded, and the exact contents Google read. Crawl stats sits on the same Settings screen, and the reply in the thread points to both.
  • Audit the file token by token instead of searching it for the string “Googlebot”. Check each variant in the table above against your own rules, including any catch-all group.
  • Leave the variants out of the file unless you want to treat them differently. A single allow group for Googlebot already covers the variants that fall back to it, and every extra named group is another chance to disallow something by accident.
  • Handle a request flood at the server or WAF, or block the one offending bot by name. Robots.txt is not a rate limit, and a rushed rewrite of the whole file is what stopped Google here.
  • Track crawl stats and indexed page counts first when access is restored. Clicks move last, and on a thin page set they may not move much at all.