Mueller: AI crawlers only use sitemaps they find on their own

anything without a submission channel cannot find the…

Summary

A sitemap at an unusual path works only for the search engines it was submitted to, and AI crawlers have nowhere to submit one. Google's Search Relations lead points to a default sitemap.xml or a linked RSS feed as what reaches them.

Sites that want their content in AI systems should keep both standard and live, and should not expect llms.txt to stand in for either.

AI training crawlers usually have nowhere to submit a sitemap. They have no equivalent of Search Console, so the only sitemaps they can use are the ones they find on their own. Google’s John Mueller said as much on the October 1 episode of the Search Off the Record podcast, titled “Do sitemaps still matter?”, and added that his own server logs show AI crawlers fetching both his sitemap file and his RSS feeds, according to Search Engine Journal’s write-up.

The advice cuts against a trick for keeping a sitemap private. Mueller laid it out: a site can give the file an unusual name, leave it out of robots.txt, and submit it to Google directly. It works for Google. Bing would probably need its own submission, Mueller said, and anything without a submission channel cannot find the file at all. What that last group loses is the whole URL list.

Two things make a sitemap discoverable with no submission anywhere. One is the default name. “Either you stick to the generic naming, call it sitemap.xml, or you focus on RSS feeds,” Mueller said, and feeds are the easier of the two to find because they are usually linked from a page’s HTML head. The other is the robots.txt Sitemap line. Under the sitemaps protocol that line is independent of the user-agent groups around it, so a disallow aimed at GPTBot does not hide the Sitemap directive from GPTBot. Robots.txt rewrites meant to fend off AI bots have taken sites out of Google before, but losing the sitemap pointer is not one of the failure modes.

Mueller was talking about training crawlers, the category GPTBot and ClaudeBot belong to. He did not name which ones he saw, and said he does not know whether AI companies document the behavior or what they do with the files. That gap matters more than the fetches do. A sitemap request in a log proves a file was read, not that a single URL in it was crawled, trained on, or cited. Attributing the request is getting harder too, since OpenAI’s always-on agents browse with no published user agent.

None of this makes llms.txt a replacement. Asked whether the Markdown file could stand in for an XML sitemap, Mueller compared it to an HTML sitemap and said Google’s systems cannot use it as a sitemap because it lacks the strict format. “I think the hope is bigger than the reality,” he said, allowing that search systems might read Markdown files someday but adding that “currently none of this happens.” In August he reported that the only crawlers on his test sites claiming to accept Markdown were SEO tools. Publishing llms.txt is cheap and harmless. Publishing it in place of a sitemap trades a format crawlers already parse for one nobody has committed to parsing.

A correctly named sitemap at the root can still go unfetched. Martin Splitt, Mueller’s Search Relations colleague, closed the episode by asking why Search Console reports “Couldn’t fetch” for a valid, public sitemap that robots.txt links. Both causes Mueller gave sit outside the file. Google’s systems may be too busy with the host to fetch it, or crawl demand is low, which Mueller tied to “the perceived quality of a website” rather than anything technical. Standard naming is the floor, not a promise of attention.

What to do

  • Keep the sitemap at /sitemap.xml and keep the Sitemap: line in robots.txt. The obscure-name route is for sites that genuinely want Google-only discovery and accept that everyone else loses the file.
  • Keep RSS working and linked from the head of every page, since that is the file AI crawlers reach most easily: <link rel="alternate" type="application/rss+xml" href="https://example.com/feed.xml">
  • Grep access logs for requests to the sitemap and feed paths from non-search user agents to see which AI crawlers pull them on your own site. Expect requests you cannot attribute to any documented bot.
  • Treat llms.txt as an experiment that sits alongside the sitemap. Nothing in Google’s current guidance gives it a job a sitemap was doing.