Back to all articles
Building in PublicJune 22, 2026
A
Latent Pulse / Building in Public
L8EntSpace

One AI Probe Will Lie to You

SYS_RENDER_OK|NODE_1856

Ask an AI engine the same question twice and you can get two different answers. So a single probe is a coin flip, not a measurement. Here are the three changes we made so our numbers admit their own uncertainty instead of hiding it.

A single AI answer is a coin flip, not a measurement

Ask an AI engine the same question twice and you can get two different answers. That is not a bug, it is how these models work (they sample). So if you probe your brand once and see "cited on 2 of 7 questions," the honest version of that number is not a clean 29%. It is "somewhere in a wide range, and we are not sure yet." Treating one probe as truth is the most common way AI-visibility dashboards mislead people. Here is what we changed so ours does not.

The problem with one pull

Imagine judging a coin as biased after three flips. Two heads and you declare it weighted, when you have actually learned almost nothing. AI probes are the same: a handful of stochastic answers cannot pin down a real rate. The number looks precise (a tidy "29%") but the uncertainty around it is enormous. Reporting the tidy number alone is a quiet lie of false precision.

Three honest fixes

We made three changes, all aimed at not fooling you (or ourselves).

First, repeat sampling. On paid plans the probe now asks each question several times per engine and takes the majority, so one weird answer does not swing the result. More samples, tighter estimate.

Second, confidence intervals. Every citation rate ships with a 95% Wilson interval (what that is), a standard way to say "the true rate is probably in this band." At small sample sizes the band is wide on purpose. If your rate is "14% (interval 8% to 64%)," movement inside that band is noise, not progress. Seeing the band stops you celebrating randomness.

Third, a real significance test on head-to-heads. When you compare against a competitor, we run a two-proportion z-test and report a plain verdict: ahead, behind, or "inconclusive at this sample size," with a p-value. No more reading a win into two bars that are basically tied.

Why this is the whole point

The entire value of a GEO measurement tool is that you can trust the number enough to spend money against it. A tool that hands you confident-looking figures with hidden uncertainty is worse than no tool, because it sends you chasing noise. Showing the interval, sampling more, and testing differences are not academic niceties. They are the difference between a dashboard you can bet on and a slot machine with a nice font.

We would rather tell you "we cannot call this yet, run more questions" than pretend a coin flip is a trend. That honesty is the product.

What to do next

When you run a probe, look at the interval before the headline number. If it is wide, add more target questions and re-sample before drawing conclusions. For competitor comparisons, read the verdict, not the bar heights. And if your rate is zero on some engines in the first place, that is usually a pathway issue, explained in Parametric vs Grounded.

Key Takeaways

  • AI answers are sampled, so one probe is a coin flip, not a measurement.
  • Paid probes now repeat-sample each question per engine and take the majority.
  • Every rate ships with a 95% Wilson confidence interval; movement inside a wide band is noise.
  • Competitor comparisons use a two-proportion z-test with an honest "ahead, behind, or inconclusive" verdict.
  • If the interval is wide, gather more questions before acting.

Ready to dominate AI search?

Start extracting high-entropy facts and tracking your Share of Voice today.

Get Your Free Report

Keep reading