Crawling & Indexing
Robots.txt, sitemaps, crawl budget, and indexing issues.
- News
Google documents how to tell Googlebot when to retry a 503
Google's reduce-crawl-rate guide now shows Retry-After examples in seconds and absolute UTC form, plus the 1-2 day limit on returning those error responses.
- News
A resolver without the new DNS root key breaks sites on Oct 11
The DNS root switches to a new signing key, KSK-2024, on October 11. Two sentinel queries defined by RFC 8509 show whether a resolver trusts it before the switch.
- News
Google puts numbers on site moves and core update recovery
The ranges reported from Gary Illyes' Barcelona session cover discovery, refresh, canonicalization and manual actions, with no sample size or definition of typical.
- News
A robots.txt rewrite to stop AI bots took a site out of Google
The rewrite allowed Googlebot but named Googlebot-Image and Google-InspectionTool, and Google's crawler docs show why robots.txt needs a per-variant audit.
- News
JavaScript error pages can make Google index another site instead
A Reddit user saw Google pick a casino page as canonical for their pages. Mueller says an indexed JavaScript error page may explain it and suggests hourly checks.
- News
Google says AdSense crawler rules also affect Ad Manager
Google's crawler docs now say Mediapartners-Google rules affect several ad products, and the crawler ignores the * group, so only a named robots.txt group stops it.
- News
Google crawled a news site less as its soft 404 count climbed
A Search Engine Land case study traces a 90% traffic loss at a crypto news site to a botched migration, then finds soft 404s growing on its other country sites.
- News
Google deindexed a 1,800-post blog that passed technical checks
A gaming blog's 1,015 pages sat at Crawled - currently not indexed with blank canonicals, and a Screaming Frog crawl found exploit scripts and download gate code.
- News
Fix soft 404s on empty category pages by how long stock is gone
Search Console flags empty category pages as soft 404s. Added content, a 302 or noindex each fits a different stockout length, and a canonical fits none.
- News
Search Console flags failed page resources that logs show loaded
A WooCommerce owner saw URL Inspection mark images and fonts as Other error while logs showed 200 OK. Replies point to the live test's limits and no-cache headers.
- News
Check server logs before blocking millions of news tag pages
An r/TechSEO thread on a news site with millions of tag and author pages weighs noindex against robots.txt disallow, and puts a server log check before both.
- News
OpenAI's search crawler now outpaces its training bot in logs
Botify's analysis of about 7 billion log events finds OAI-SearchBot up 3.5x since GPT-5, and OpenAI's docs show GPTBot blocks do not cover ChatGPT search.
- News
ChatGPT reportedly gets Google results from a scraping service
Botify says OpenAI, Meta and Perplexity reportedly buy scraped Google results from SerpAPI, and explains the jobs of GPTBot, OAI-SearchBot and ChatGPT-User.
- News
A noindex tag does not stop AI bots from using your bandwidth
Botify ties rising AI crawler traffic to bandwidth and uptime costs. Google's docs show why noindex cannot cut that load: bots must fetch the page to read it.
- News
Mueller can't confirm splitting sitemaps by age changes crawling
Mueller gave five reasons sites split XML sitemaps. File limits and hreflang size justify a split; splitting by content age is a theory he could not confirm.
- News
Blocking CSS and JS in robots.txt breaks how Google renders pages
An r/TechSEO poster saw 79% of crawl requests go to CSS and JS files. The replies say keep them crawlable, as Google's docs do, and find out why they are refetched.
- News
A wildcard DNS record got made-up subdomains indexed and ranking
A dating network's wildcard DNS left gibberish subdomains indexed and drawing traffic. Google's canonical docs back a path-preserving 301 to the www host.
- News
Indexing API got ordinary pages indexed fast in a Reddit test
A site owner's split test got 94% of Indexing API pages indexed in a week, against 8.4% by sitemap. Google supports the API only for job and livestream pages.
- News
Mueller: odd canonical picks often trace to what Googlebot saw
John Mueller lists why Google treats URLs as duplicates, from bot challenge pages and unrendered JavaScript shells to mobile comparison and parameter guessing.