AI Crawler Management
Blocking training bots, allowing retrieval bots, AI-specific robots.txt.
- News
Mueller: AI crawlers only use sitemaps they find on their own
Google's John Mueller says AI training crawlers have no console for sitemap submission, and that his own server logs show them fetching his sitemap and RSS files.
- News
OpenAI's always-on agents browse with no published user agent
OpenAI's dots and Google's AI Mode monitoring both keep re-checking the web after the user leaves, and OpenAI's bot docs name no user agent for a dot's browser.
- News
Cloudflare tests charging AI agents instead of blocking them
Cloudflare's Monetization Gateway lets a site price agent requests and answer with HTTP 402, but it is a US-only closed beta, so the job today is measuring bots.
- News
Cloudflare's AI training block now stops Googlebot too
Since September 15, Cloudflare's Training block also shuts out Googlebot, Bingbot and Applebot. Disallow AI Training refuses training and keeps search.
- News
Robots.txt only stops AI crawlers that choose to obey it
Reverse DNS for Googlebot, OpenAI's published IP lists, and CDN rate limits deal with the AI bots a Reddit thread found hitting parameter URLs despite robots.txt.
- News
A WordPress host silently blocked AI training bots on one site
Logs and curl tests in Search Engine Land show WP Engine refusing ClaudeBot, GPTBot and Amazonbot while PerplexityBot got through, plus a test for your host.
- News
Google signs some AI agent requests so sites can spot fakes
Google's experimental Web Bot Auth signs some Google-Agent requests with keys published at agent.bot.goog, yet Google still says to keep IP and reverse DNS checks.
- News
Cloudflare can redirect AI training bots to your canonical URLs
The AI Crawl Control toggle sends GPTBot and ClaudeBot a 301 to same-origin canonicals, so every page canonicalized elsewhere drops out of their crawls.