A WordPress host silently blocked AI training bots on one site

robots.txt line: Disallow: /

Summary

A managed WordPress host can refuse AI crawlers from a layer no site owner can see or configure, and the errors look like ordinary rate limiting. On one WP Engine site, training crawlers were refused while PerplexityBot and ChatGPT-User got through.

Request your own pages as each AI bot to learn what your host does. Then write the policy you actually want into robots.txt: search bots allowed, GPTBot and ClaudeBot disallowed.

WP Engine, the managed WordPress host, was returning HTTP 429 (“too many requests”) errors to ClaudeBot, GPTBot and Amazonbot on searchinfluence.com, according to Will Scott’s investigation in Search Engine Land, published on 2026-05-06. The block ran inside the host’s own platform, between Cloudflare and WordPress. None of the site’s security tools showed it, and no setting in the WP Engine portal controlled it.

Scott, CEO of the marketing agency Search Influence, was looking at his own firm’s site because of a citation gap. Scrunch, the AI citation monitoring tool the team uses, showed the site at 37.8% presence in Google AI Mode and 0% in Claude over 30 days. Seven days of Cloudflare logs then showed ClaudeBot and GPTBot each rate-limited on 29% of requests. ChatGPT-User and PerplexityBot, which fetch pages during live user queries, were rate-limited on none.

How the block was found

The team ruled out a WordPress security plugin, that plugin’s firewall logs, and a subscription to Sucuri’s cloud firewall that turned out never to be in the request path. Cloudflare’s security events showed zero actions against ClaudeBot on a day that produced 608 ClaudeBot 429s. The deciding test was plain: 60 fast curl requests with a ClaudeBot user agent all returned 429, and the same paths with a browser or Googlebot user agent all returned 200. The block keyed on the user agent, not the path or the request rate. A curl -I then showed x-powered-by: WP Engine.

The same test run against other AI user agents gave these results:

User agentResult
ClaudeBot60 of 60 returned 429
GPTBot8 of 10 returned 429, 2 served from cache
Amazonbot10 of 10 returned 429
Bytespider10 of 10 returned 520
anthropic-ai (older Anthropic agent)10 of 10 returned 200
CCBot (Common Crawl)10 of 10 returned 200

Two details hide the block from anyone looking at dashboards. Pages already in WP Engine’s edge cache were served to ClaudeBot normally, so the same bot got 1,054 successful responses and 608 errors in one day. The 429 status also reads as a rate limit in every firewall tool, which sent the team hunting for rate-limit rules at the wrong layer. Scott quotes WP Engine’s security documentation declining to give further information about its firewall.

What the block does not explain

As a training opt-out, the host’s list leaks. It lets through CCBot, whose crawls end up in many training datasets, and Anthropic’s older anthropic-ai agent. Scott’s reading is that the list targets the training crawlers of mid-2024.

The Claude citation gap probably needs another explanation. Anthropic’s help page describes three bots: ClaudeBot collects training data, Claude-SearchBot indexes pages for search, and Claude-User fetches a page when a person asks Claude something. Blocking ClaudeBot tells Anthropic to leave the site out of training. By Anthropic’s own description, ClaudeBot is not the bot that decides whether Claude can cite you. Scott’s published curl results do not include Claude-SearchBot or Claude-User, so the host block and the zero in Claude are not linked by evidence yet.

Some of the blocked traffic was not Anthropic at all. Nearly all of the “ClaudeBot” requests in the 24-hour sample came from one Microsoft Azure IP address rather than Anthropic’s published ranges, which Scott calls almost certainly a scraper faking the user agent. Refusing that traffic is correct, and it is the same problem covered in an earlier article on AI crawlers that fake their identity.

What to do

A host can make bot decisions for you below robots.txt, and the only way to see them is to request your own pages as each bot. OpenAI says each of its bots is controlled separately, and Perplexity says PerplexityBot is not used to train models, so a site can let search and retrieval bots in and keep training crawlers out.

  1. Test the bots that feed AI answers first, since those cost you citations. Scott’s curl table did not include OAI-SearchBot or Claude-SearchBot. Use a browser user agent as the control, and pick pages that are unlikely to be cached.
for bot in OAI-SearchBot PerplexityBot Claude-SearchBot ChatGPT-User Claude-User GPTBot ClaudeBot; do
  printf "%s " "$bot"
  curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 (compatible; $bot/1.0)" "https://example.com/some-page/"
done
curl -sI -A "Mozilla/5.0 (compatible; ClaudeBot/1.0)" "https://example.com/some-page/" | grep -iE "x-powered-by|x-cache"
  1. Write the intended policy into robots.txt, even if the host already blocks the training crawlers. Anthropic says robots.txt is the way to opt out and that IP blocking may not hold, because it stops ClaudeBot from reading the file. OpenAI’s crawler documentation says robots.txt changes take about 24 hours to reach its systems.
User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /
  1. If a search bot fails the test and your host offers no setting for it, open a support ticket and ask what rule fires. If you run your own firewall, Perplexity’s crawler documentation gives an allow rule that matches the user agent and checks the request against its published IP ranges, which also keeps faked user agents out.