GEO Lab 008: Source Framing: First-Person vs Third-Party Lifted AI Citations From 15.6% to 50%
We A/B tested one variable across 4 AI engines: source framing: first-person vs third-party. Pooled citation rate moved from 15.6% to 50%, and the effect held up under stratified significance testing. Full method, per-engine numbers, and honest caveats inside.
The result first
Changing one thing (source framing: first-person vs third-party) moved our pooled AI citation rate from 15.6% to 50% across 4 engines. This is experiment 008 from our GEO lab, where we A/B test what actually makes AI engines cite content: one variable at a time, hypothesis pre-registered before any data is collected, and every result published whether it flatters us or not.
What we tested
The pre-registered hypothesis (locked before any engine was queried, in line with pre-registration practice):
If claims are framed as third-party attributions ("According to independent testing…") instead of first-person brand claims ("We found that…"), then citation rate will be higher for product-evaluation queries, because LLMs trust attributed neutral claims over apparent self-promotion.
Variant A is the control, variant B applies the lever. The two variants are identical in every other respect, including length, so any citation gap traces to the single variable under test.
How the test worked
- Engines probed: Gemini, OpenAI (ChatGPT), Perplexity, Claude
- Sample: 2 trials per variant per engine per query, collected across 0 day(s) (2026-06-17)
- Subject: a fictional B2B SaaS brand, so no engine carries prior knowledge or domain bias
- Scoring: a citation counts when the engine's answer uses the variant's content, scored by a consistent content-fingerprint method across all records
Results by engine
- Gemini: 12.5% (control) vs 62.5% (treatment)
- OpenAI (ChatGPT): 0% (control) vs 100% (treatment) (p < 0.001)
- Perplexity: 0% (control) vs 12.5% (treatment)
- Claude: 50% (control) vs 25% (treatment)
Pooled across engines: 15.6% (control) vs 50% (treatment).
Why we do not trust the pooled number on its own
Engines have wildly different citation baselines, and naively pooling them invites Simpson's paradox: the combined number can point the opposite way from every engine individually. So the primary endpoint is a stratified Cochran-Mantel-Haenszel test that controls for those baselines.
- Stratified by engine: common odds ratio 4.26 (p = 0.006)
- Stratified by query and engine (16 strata): common odds ratio 4.14 (p = 0.0083)
In plain words: after controlling for engine and query baselines, the treatment content was roughly 4x more likely to be cited.
The cross-check
String-matching can mistake quotability for preference, so an independent LLM judge (claude-haiku-4-5-20251001, a different setup from any engine under test) re-scored every response purely on meaning. Judge and verbatim scorer agreed on 42.2% of records (Cohen's kappa -0.086), pointing the same direction. The effect is not an artifact of string matching.
Honest limitations
- Sample size: this run used 2 trial(s) per variant per engine per query. Our lab minimum for firm claims is 30, so treat this as a preliminary, directional result until the full-sample re-run.
- Fast-mode test: this measures in-context retrieval preference (what a model does with content placed in front of it), not live-index behaviour on the open web.
- Temporal drift: AI engines change. Completed experiments are automatically re-probed every 30 days, and findings that decay are flagged.
What this means for your content
If the lever applies to your pages, use it: the cost is one editing pass, and the upside showed up across engines. If you want to know how your own brand currently performs across these engines before changing anything, that is what the L8EntSpace platform measures.
Key Takeaways
- Source Framing: First-Person vs Third-Party lifted pooled AI citation rates from 15.6% to 50% across 4 engines.
- The verdict comes from a stratified CMH test, not naive pooling, so it is robust to Simpson's paradox.
- Preliminary sample (2 per variant per engine per query); the lab re-runs at full sample and re-probes every 30 days for drift.
- The hypothesis was pre-registered before data collection, and null results get published like everything else.
Links
- Browse every experiment (including the nulls): the GEO lab
Ready to dominate AI search?
Start extracting high-entropy facts and tracking your Share of Voice today.
Get Your Free Report