Test chatbots with web search off to see where brand errors start

JSON-LD line: "sameAs": [

Summary

When a chatbot gets a brand wrong, the error can come from what the model learned in training, from the pages it fetches now, or from answers that never need the brand. Each fault has a different fix, so find which one applies first.

Ask a chatbot about the brand with web search off, then log what it cites with search on. Clean up name variants, sameAs links and stale third-party profiles before writing new content.

A Search Engine Land piece by organic marketing consultant Jes Scholz, published April 30, 2026, argues that when ChatGPT or an AI Overview gets a brand wrong, the fix depends on where the error started. Scholz splits AI brand visibility into three layers and says the same tactics do not work on all of them. The useful part of the piece is that each layer comes with its own test, so a brand can find the fault before spending money on content.

Three places a brand description can break

Scholz gives each layer a test:

  • Training. The brand’s historical footprint: press, reviews, documentation and old forum threads the model absorbed. Scholz’s test is to ask a chatbot to describe the brand with web search turned off.
  • Retrieval. The indexed pages, product feeds and APIs an AI system can fetch at answer time and cite. Scholz says this is where crawling, indexing and rendering matter most. The test is to run branded and category prompts daily in an LLM tracker and note which sources keep getting cited.
  • Generation. The finished answer in AI Overviews, AI Mode or ChatGPT. Scholz says a brand gets written into an answer only when the answer needs it. The test uses the same tracker data, read for how the brand is mentioned and what it appears next to.

The search-off test matters because retrieval can hide training errors. Systems built on retrieval-augmented generation add fetched documents to the prompt and are set up to prefer that fresh text over what the model learned in training. So a correct answer with search on can sit on top of an outdated picture. The likelier outcome when retrieval finds nothing useful is that the model falls back on that older picture.

Take a made-up example. A software company renamed itself two years ago. With search off, a chatbot describes it under the old name and the old product line. That is a training-layer problem, and new on-site copy will not fix it quickly. With search on, the answer cites a directory listing that still carries the old name, which is a retrieval problem fixed by editing one listing. If both descriptions are right but the brand never shows up for category prompts, the fault is at generation.

The layers also map onto different crawlers, though Scholz’s piece does not draw the link. Training crawlers such as GPTBot and ClaudeBot feed future models. Search bots such as OAI-SearchBot and PerplexityBot, along with user-triggered fetchers such as ChatGPT-User, feed retrieval. A robots.txt rule that blocks all of them together can cut the brand out of today’s citations while doing nothing about what existing models already learned.

Name variants split the signal

Scholz’s most practical point concerns naming. A brand often has a display name spaced or cased inconsistently, a legal name, a domain, an abbreviation and a legacy name. People merge these without thinking. Scholz argues models merge them only when the pattern makes it obvious: “Allow your brand to be written five different ways and split your visibility signals five times.”

The fix Scholz proposes is a canonical brand bio in the form “[Brand] is a [market category] for [audience] who need [use case], differentiated by [proof]”, plus fixed rules for spacing, casing and abbreviation. If the bio could also describe a competitor, Scholz says to rewrite it. The same wording then goes into structured data, sameAs references, industry publications, partner sites, reviews and community discussions. Pages that describe the brand differently get consolidated or removed.

The piece offers no measurements, so treat its steps as a low-cost checklist rather than tested fixes. Most of them overlap with the entity cleanup many sites need anyway.

What to do

  1. Run both tests on the brand’s main prompts: search off, then search on with the cited sources logged. Write down which layer is failing before changing anything.
  2. Write the canonical bio and the naming rules, and check the bio against your closest competitor.
  3. Mark up the organization with Organization schema and list every official profile in sameAs. Scholz’s piece does not mention alternateName or legalName, but they are the natural place to declare the legal and legacy names you cannot erase, so those names point to the same entity.
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "Example Analytics",
  "legalName": "Example Analytics Ltd",
  "alternateName": ["ExampleAnalytics", "OldName Data"],
  "url": "https://www.example.com/",
  "description": "Example Analytics is a real-time dashboard tool for enterprise data teams who need governed reporting.",
  "sameAs": [
    "https://www.linkedin.com/company/example-analytics",
    "https://github.com/example-analytics"
  ]
}
  1. Fix the third-party pages that carry old names or descriptions, starting with the ones your tracker shows being cited.
  2. State proof plainly on the page. Scholz lists awards, benchmarks, customer numbers and policies as things to make easy to quote. AI search scores single passages rather than whole pages, so each fact should make sense on its own without the paragraphs around it.