Hreflang checkers pass tags that AI crawlers never see

Diagram from Prerender.io: What hreflang checker misses on JavaScript websites: What href…
Diagram: Prerender.io, "What hreflang checker misses on JavaScript websites"

Summary

A clean hreflang report says nothing about what crawlers received. Audit tools execute JavaScript before reading tags. Googlebot's first pass and the AI crawlers do not.

The check that settles it takes two minutes: disable JavaScript and view source on one page per locale, confirming that the full hreflang set and the canonical are both in the served HTML.

A hreflang checker runs the page’s JavaScript before it looks for tags, so a site can score zero errors on an audit and still have served no language signals to anyone. Prerender.io’s post on what hreflang checkers miss on JavaScript sites, updated September 30, spells out the sequence: the tool requests each URL, lets the server respond, executes client-side JavaScript, then scans the DOM to confirm that every language variant is declared, that return links point back, and that canonicals line up. Anything the framework injected counts as present. Prerender.io sells a prerendering service and has an interest in that conclusion, but the mechanism is checkable on any locale page in a couple of minutes.

Googlebot reads a page in two passes, which the post calls Wave 1 and Wave 2, and only the second runs JavaScript. The first takes the raw HTML off the server and processes the links and meta tags in it right away. The rendered version arrives later, out of a render queue shared across the whole web, and Prerender.io’s post puts the practical gap at anywhere from a few hours to several weeks depending on site size and crawl priority. Language targeting therefore starts from the document that may carry no hreflang at all, while the audit grades the one that does.

A delay would be survivable if hreflang were a per-page signal. It is not. Google’s hreflang documentation, quoted in the post, requires that if page X points to page Y then page Y points back to page X, and says the annotations may be ignored or misinterpreted where that reciprocity is missing across the pages using hreflang. One locale page whose tags arrive only after rendering can leave the entire group unconfirmed. The post walks through a case where a US page lists every alternate in raw HTML while the German and French pages inject theirs with JavaScript: on the first pass over those alternates, the return links to the US page do not exist, so the relationship stays unverified even though the US page itself is flawless. A rendered-DOM audit reads all three after render and reports nothing wrong.

Canonicals break along the same seam, because many frameworks inject the canonical in the same head update as the hreflang tags. When the raw HTML ships a default canonical or none and JavaScript rewrites it afterwards, Google receives two answers for the same URL at two different times, and each annotation is supposed to name the canonical version of its locale. When Google cannot tell which URLs belong in the cluster, the post argues, ignoring the hreflang signals is the safest option from its side. The same late-head problem produced Google indexing Next.js pages with an empty title and no canonical, so the failure is not specific to one framework.

Render priority is also uneven inside a single site, and the post notes that high-traffic pages tend to get rendered sooner. On a large multi-language catalogue that means some members of a cluster show their tags while their siblings sit in the queue for weeks, from identical code. Spot-checking two URLs will not find it.

The AI half does not resolve on its own

Prerender.io’s post states flatly that GPTBot, ClaudeBot, PerplexityBot and the other bots feeding AI search surfaces do not execute JavaScript at all. They read raw HTML and stop, with no second pass, so a client-side hreflang set leaves every AI-driven search surface working with zero language signal. For those bots the render-queue delay is not a delay. The tags never arrive.

Search Console will not name the problem either. Google removed the dedicated International Targeting report in September 2022, so there is no direct hreflang diagnostic left. What surfaces instead is in the Pages report, where Google picks the wrong country’s page as canonical and the intended localized URL sits unindexed or misrouted.

What to do

Audit the raw HTML, not the rendered page, and do it per locale:

  • Disable JavaScript, open view-source on one page per locale, and search it for hreflang. Google dropped its advice to test pages with JavaScript turned off, but for hreflang the served HTML is still the document the first pass and every AI crawler read. At scale, curl -s https://example.de/ | grep -iE 'hreflang|canonical' gives the same answer per URL.
  • Check the alternates, not just the page whose targeting looks broken. The missing return link is usually on the other side of the cluster.
  • Confirm the canonical is present in the served HTML and points where you expect, since an absent or default canonical is enough on its own to get the annotations dropped.
  • If the tags exist only after render, move the head into the server response through server-side rendering, static output, or prerendering. Prerender.io recommends its own product; any path that ships the complete head in the first response closes the same gap.