Two AI visibility tools reported 41% and 23% for one client

Summary

Two AI visibility platforms reported 41% and 23% presence for the same client. Mike King's argument is that neither is wrong, because AI answers are probabilistic and personalized, so there is no true number to hit.

The fix is precision, not accuracy. Pin the prompt set, model, mode and run count, report movement inside one tool's own band, and tie it to referral data.

What happened

Two AI visibility platforms reported 41% and 23% presence for the same client. iPullRank founder Mike King opened a September 17 piece with that gap. His argument is that neither figure is wrong, because there is no true number for either one to miss.

The distinction King borrows from measurement science is simple. “Accuracy is how close a measurement is to the true value. Precision is how close repeated measurements are to each other.” AI answers are generated probabilistically and personalized per user. There is no single response for a tracker to match against.

The variance underneath those percentages is large. King cites Mike Sonders’ research in Search Engine Land, which ran 12 buyer-intent prompts 100 times each through logged-out ChatGPT. Across 100 runs of one prompt, about 44 different brands appeared, roughly 10 per response. Only about five brands, 11% of the pool, showed up in 80% or more of responses. 72% appeared in fewer than one response in five.

Why it matters

The board deck problem is a sampling problem. King publishes the confidence interval math for a single prompt at a true 50% appearance rate. Five runs give a margin of ±44 points. Ten runs narrow it to ±31, 20 runs to ±22, 50 runs to ±14, and 100 runs to ±10. A tracker that runs each prompt once per collection cannot separate a real gain from a dice roll. King’s own phrasing is that appearing for a given prompt on a given day “is closer to a dice roll than a ranking.”

Two more methodology forks widen the gap between vendors. A Semrush study from June compared ChatGPT’s Instant and Thinking modes across 100 prompts and found roughly 25% domain overlap between them. Thinking mode issued more than four times as many fan-out queries and cited nearly twice as many sources. It also pulled from 99 domains that never appeared in Instant mode.

Prompt phrasing moves the number as well. An SSRN working paper from Ehrlinspiel, Landwehr and Rudzki analyzed about 38,000 responses across five engines. Ranking-style prompts surfaced up to 20% more brands than open-ended ones. A vendor whose prompt library leans on “best X” questions will report a higher presence score than one that asks open questions, on an identical account.

Disagreement between measurements of AI search is now the pattern rather than the exception. Six studies of AI traffic conversion published this year reached conflicting conclusions on the same basic question. Vendors keep shipping presence metrics regardless, including Botify’s AI Visibility product, which replaces rank with mentions.

What to do

Freeze the instrument before the next reporting cycle. Pin the model, the mode, the account state, the geography, the prompt set and the number of runs per prompt. Write that configuration down. Any change to those inputs starts a new trend line.

Report movement inside one tool’s own band. Pick the platform you intend to keep and stop quoting the competitor’s number in the same deck. The two are not measuring the same system.

Test precision yourself before you trust a trend. Run the same prompt set twice under identical conditions and compare the two results. If they disagree by more than the movement you plan to report, the movement is noise.

Ask your vendor a short list of questions that King spells out:

  • Which model and mode are pinned, and how are version changes announced.
  • How many runs per prompt per collection, and does the dashboard show a range rather than one number.
  • Where does collection originate geographically, and can you control it.
  • Is there a methodology changelog, and can you export raw responses rather than extracted metrics.

Couple the visibility number to referral data before you act on it. King’s framing is that “visibility metrics are channel metrics, and they’re proxies. Performance metrics are the actuals.” If presence moves and sessions from AI assistants do not, the move may be noise.

Watch out for

API and interface data do not mix. API collection measures controlled model behavior. Scraping the consumer interface measures what users actually see, including product-layer features the API never runs. Two vendors using different collection methods are reporting on two different systems. Google AI Overviews has no developer API, so a vendor claiming API collection there is approximating or scraping.

Logged-in collection contaminates itself. ChatGPT’s Memory feature lets a collection account learn from its own prompt history. The instrument changes what it measures over time. Logged-out collection avoids that but misses the personalization most real users get.

“Mention” is not a standard definition. A brand name in body text, a linked citation, and a recommendation in a ranked list are three different events. Tools that count them differently produce different percentages from identical responses.