How We Tried to Break Our Own GEO Findings (and What Survived)
Before we show a finding to anyone, we try to break it. Any weakness we find ourselves is one a skeptical client cannot use against us later. Here are the confounds we hunted in our own experiments, and what held up under the tougher test.
The fastest way to lose trust is to be caught overclaiming
So we attack our own results first. Before we show a GEO finding to anyone, we try to break it. The logic is simple: any weakness we find ourselves is one a skeptical client (or a rival) cannot use against us later. Recently we ran a batch of hardening checks against our own experiments to see which findings would survive a tougher test. Most did. Here is what we checked and why it matters to you.
The confounds we hunted
A few specific things can quietly inflate an AI experiment.
Generation settings. If one engine is allowed to write longer answers than another, the longer answers cite more sources, and you mistake answer length for a real effect. We pinned the temperature and capped output length the same way across engines, and we log those settings on every trial.
The self-judging problem. Our experiments include Claude, so judging answers with Claude invites bias toward Claude. We switched the neutral judge to a different model family and recorded which family did the judging.
Counting agreement honestly. We used to report raw percent agreement between two scoring methods. Raw agreement is flattering, because most answers are "not cited" and both methods agree on those by default. We switched to chance-corrected scores (Cohen's kappa and Gwet's AC1), which strip out the easy agreement.
The big one: trials that are not independent
This is the subtle one. If you ask the same question several times, those answers are not independent of each other (they share the same wording and intent). Our main test used to group results by engine only. We changed it to group by question and engine together, so repeated answers to the same question stay in their own bucket rather than counting as fresh evidence. We kept the looser test alongside as a sensitivity check, and if the two disagree, we flag the result inconclusive rather than pick the flattering one.
We also added a power note to every finding (statistical power is the chance of detecting a real effect). When a result is null ("no detectable effect"), we now state the smallest effect we could have caught at that sample size. That turns an ambiguous null into an honest one: "no effect larger than X," rather than the lazy "it does not matter."
What survived
We re-ran every experiment under the stricter rules. Reassuringly, none of our significant findings flipped to null under the harder test, and the nulls stayed null. That is the outcome you want: results that do not depend on the generous version of the maths.
What to do next
When you evaluate any GEO claim (ours or anyone's), ask the boring questions: were the generation settings held constant, who judged the answers, and were repeated queries treated as independent? If a vendor cannot answer those, the number is decorative. For the programme-wide layer that sits on top of all this, read Run Enough Experiments and One Will Look Significant by Luck.
Key Takeaways
- We pin temperature and cap output length across engines, and log both, to remove a length confound.
- The neutral judge is never from the same family as the engine under test.
- Agreement is reported chance-corrected (Cohen's kappa, Gwet's AC1), not raw percent.
- The main test groups by question and engine, so repeated answers are not counted as independent evidence.
- Re-running every experiment under the stricter rules flipped none of our significant findings.
Ready to dominate AI search?
Start extracting high-entropy facts and tracking your Share of Voice today.
Get Your Free Report