Back to all articles
Lab NotesJune 17, 2026
A
Latent Pulse / Lab Notes
L8EntSpace

GEO Lab 017: Opening Question vs Opening Statement: No Significant Effect (a Null Result Worth Knowing)

SYS_RENDER_OK|NODE_1856

We A/B tested one variable across 4 AI engines: opening question vs opening statement. The result was a null: no statistically significant effect. We publish those too, because knowing what does not move AI citations is as valuable as knowing what does.

The result first

Nothing happened, and that is the finding. This is experiment 017 from our GEO lab, where we A/B test what actually makes AI engines cite content. We changed one thing (opening question vs opening statement) and the citation rates (37.5% vs 25% pooled) were not statistically distinguishable. Publishing null results keeps the wins honest: a lab that only reports positives is a marketing department.

What we tested

The pre-registered hypothesis (locked before any engine was queried, in line with pre-registration practice):

If content opens by restating the query as a direct question instead of opening with the answer, then citation rate will be lower for how-to and definitional queries, because LLMs prefer answer-mode sources that present extractable claims over question-mode sources that defer the answer.

Variant A is the control, variant B applies the lever. The two variants are identical in every other respect, including length, so any citation gap traces to the single variable under test.

How the test worked

  • Engines probed: Gemini, OpenAI (ChatGPT), Perplexity, Claude
  • Sample: 2 trials per variant per engine per query, collected across 0 day(s) (2026-06-17)
  • Subject: a fictional B2B SaaS brand, so no engine carries prior knowledge or domain bias
  • Scoring: a citation counts when the engine's answer uses the variant's content, scored by a consistent content-fingerprint method across all records

Results by engine

  • Gemini: 62.5% (control) vs 25% (treatment)
  • OpenAI (ChatGPT): 50% (control) vs 25% (treatment)
  • Perplexity: 25% (control) vs 25% (treatment)
  • Claude: 12.5% (control) vs 25% (treatment)

Pooled across engines: 37.5% (control) vs 25% (treatment).

Why we do not trust the pooled number on its own

Engines have wildly different citation baselines, and naively pooling them invites Simpson's paradox: the combined number can point the opposite way from every engine individually. So the primary endpoint is a stratified Cochran-Mantel-Haenszel test that controls for those baselines.

  • Stratified by engine: common odds ratio 0.56 (p = 0.4227)
  • Stratified by query and engine (16 strata): common odds ratio 0.67 (p = 0.4334)

Neither stratification produced a significant effect, which is why we call this a null result.

The cross-check

String-matching can mistake quotability for preference, so an independent LLM judge (claude-haiku-4-5-20251001, a different setup from any engine under test) re-scored every response purely on meaning. Judge and verbatim scorer agreed on 43.8% of records (Cohen's kappa 0.106), pointing the same direction. The effect is not an artifact of string matching.

Honest limitations

  • Sample size: this run used 2 trial(s) per variant per engine per query. Our lab minimum for firm claims is 30, so treat this as a preliminary, directional result until the full-sample re-run.
  • Fast-mode test: this measures in-context retrieval preference (what a model does with content placed in front of it), not live-index behaviour on the open web.
  • Temporal drift: AI engines change. Completed experiments are automatically re-probed every 30 days, and findings that decay are flagged.

What this means for your content

Skip this lever. Effort spent here is effort not spent on changes with demonstrated effects. Our other experiments in the lab cover levers that did move citations.

Key Takeaways

  • Opening Question vs Opening Statement produced no statistically significant change in AI citation rates (37.5% vs 25% pooled).
  • The verdict comes from a stratified CMH test, not naive pooling, so it is robust to Simpson's paradox.
  • Preliminary sample (2 per variant per engine per query); the lab re-runs at full sample and re-probes every 30 days for drift.
  • The hypothesis was pre-registered before data collection, and null results get published like everything else.

Links

  • Browse every experiment (including the nulls): the GEO lab

Ready to dominate AI search?

Start extracting high-entropy facts and tracking your Share of Voice today.

Get Your Free Report

Keep reading