GEO Lab

What actually changes AI citations: tested, not guessed

Most GEO advice is confident guessing. The GEO Lab runs controlled A/B experiments instead: pre-registered hypotheses, identical content except one variable, real queries across multiple AI engines, and a proper statistical test. We publish every result here (including the preliminary ones and the failures), with the caveats stated up front. If a number on this page has an asterisk in spirit, we write the asterisk out.

⚠ Preliminary001-statistical-anchors · 2026-06-10

Do statistical anchors in the opening sentence increase AI citation?

Hypothesis

If the opening sentence contains a specific number ("cut latency 43%"), citation rate will be higher than a version with vague language ("improved latency significantly") for product-performance queries. LLMs weight precise, citable data points as credibility signals.

Setup

Engines: Gemini, ChatGPT, Perplexity, Claude. 32 trials per variant across 4 engines (8 per engine), below our pre-registered threshold of 30 per engine.

Result

21.9%

Variant A (vague) citation rate

50%

Variant B (specific number) citation rate

+62.5pp, p=0.007

Significant effect (Claude only)

The specific-number variant (B) was cited in 16 of 32 trials (50%) versus 7 of 32 (21.9%) for the vague variant (A). The only statistically significant per-engine effect was on Claude (+62.5pp, p=0.007), which was not the pre-registered primary platform. Gemini returned 8/8 null responses due to API failures (since fixed), so its data is missing entirely.

Why this verdict

Sample size (n=8 per variant per engine) is below the pre-registered n≥30; the significant effect appeared on Claude rather than the pre-registered primary platform (Perplexity); and Gemini data is missing due to API failures. Encouraging direction, but not yet a finding we will state as fact.

What we do next

Re-run at n≥30 per engine with the Gemini API fix in place, keeping Perplexity as the pre-registered primary platform. If the effect holds, it graduates to "supported" and we will say so plainly.

How the lab works

  1. Pre-register the hypothesis: the prediction, primary platform, and sample threshold are written down before the run.
  2. Change exactly one thing: variant B is identical to variant A except for the variable under test.
  3. Probe real engines: the same query set runs across Gemini, ChatGPT, Perplexity, and Claude, with randomised variant order.
  4. Test for significance: a two-proportion z-test decides whether the difference beats chance.
  5. Publish whatever comes out: supported, rejected, null, or preliminary. Preliminary means "promising but underpowered," and we say so.

Want to run experiments like these on your own content? The Citability Lab in the L8EntSpace dashboard runs the same head-to-head methodology on your drafts. Or read more on the blog.