Blocking CSS and JS in robots.txt breaks how Google renders pages
Summary
Blocking stylesheets and scripts in robots.txt stops Googlebot from rendering pages the way a browser does. Google's own robots.txt documentation keeps those files open for exactly that reason, and a large share of crawl requests going to them is usually normal.
When resource crawling looks heavy, find out why the same files get fetched again and again, starting with caching headers. Content that matters belongs in the HTML, since AI crawlers reportedly run no scripts.
A site owner on r/TechSEO found that 79% of their crawl budget went to “page resource load”, and most of those requests were for CSS and JavaScript files. The owner asked whether to disallow those files in robots.txt. The replies in the Reddit thread were close to unanimous: no, because Googlebot and other bots would no longer be able to render the pages correctly. The poster had not made the change and said they were glad they asked first.
Google’s documentation takes the same position. The example file on Google’s robots.txt specification page covers this exact case. That file blocks an /includes/ directory of .css and .js files for all crawlers, then adds an Allow for Googlebot, with a comment saying “Google needs them for rendering”:
User-agent: *
Disallow: /includes/
User-agent: Googlebot
Allow: /includes/
Rendering is the step where Googlebot runs the page’s scripts and applies its styles to see the page as a browser would. If it cannot fetch those files, Google indexes a page that is only partly built, and anything JavaScript adds to the page may be missing. Blocking the files trades a crawl statistic for an indexing problem.
A high share of resource requests is not waste in itself. Those fetches are how Google renders the site. One commenter put the useful question plainly: examine why Googlebot needs to fetch the files so often. Another suggested that, depending on the stack, better caching headers could stop Google from fetching the files more than it needs to. The likelier causes of repeated fetches are files served without long-lived caching or file URLs that change on every deploy. A site that splits its code into many small files also produces more requests than one that bundles them.
One caveat applies to the caching advice. Google’s robots.txt page mentions max-age in Cache-Control headers only in connection with how long Google caches the robots.txt file itself. It says nothing about how Googlebot treats cache headers on CSS or JS. Caching is a cheap fix worth trying, but it is not a documented crawl setting.
A separate point in the thread concerns AI crawlers. One commenter said Googlebot, and probably Apple’s crawler, execute JavaScript, but AI crawlers such as ClaudeBot appear to take the HTML and run nothing. If that holds, anything a page adds with JavaScript, including links, is invisible to those crawlers whatever robots.txt allows. That is one practitioner’s observation, not documented behavior.
What to do
- Keep CSS and JavaScript crawlable. If robots.txt blocks a directory that holds them, add an Allow rule for Googlebot, as Google’s example does.
- Open the Crawl Stats report in Search Console and check the Page resource load and By file type breakdowns. Then use server logs to see which CSS and JS URLs Googlebot requests repeatedly. The same file fetched many times points to a caching or file-naming problem, not a robots.txt problem.
- Check the Cache-Control headers on your static files and fix any that prevent caching.
- Put the content and links you need indexed in the server-sent HTML, so crawlers that do not run JavaScript still see them.