Skip to content
CiteRank AI
Research Paper

Sampling Bias in Prompt Corpora for Generative Engine Evaluation

A working paper on why scraped keyword lists produce unstable AI visibility measurements, and a stratified sampling design that makes prompt-level results reproducible across runs and engines.

CiteRank AI Research·18 June 2026·28 min readIllustrative data
MeasurementSamplingMethodologyReproducibility
Corpus size studied
40k
Modelled prompts across 6 categories
Engines compared
8
Same corpus replayed on each engine
Replays per prompt
5
Used to estimate run-to-run variance

Key takeaways

  • Prompt corpora scraped from keyword tools over-represent head terms and under-represent the comparative and advisory prompts that AI answers actually serve.
  • Run-to-run variance falls sharply once the corpus is stratified by intent, buying stage and geography rather than by search volume.
  • Reporting a visibility score without the corpus design and sample size is not a measurement — it is an anecdote.
  • Confidence intervals should be published per engine, because engine-level variance differs by more than an order of magnitude.
Resource

Working paper (v1.2)

Pre-print, not peer reviewed. Placeholder record — the PDF is issued on request while the method version is finalised.

Format
PDF, ~24 pages
Version
v1.2 — June 2026
Licence
CC BY 4.0 (sample data)
Citation
CiteRank AI Research (2026)

Problem statement

Generative engines answer questions, not queries. A corpus built from keyword exports therefore samples a different population than the one being measured, and the resulting score drifts whenever the engine reweights its retrieval.

This paper formalises that gap and proposes a stratified alternative that can be reproduced by a third party from the published design alone.

Proposed sampling design

Prompts are drawn from a frame defined by intent class, buying stage, geography and specificity, with quotas fixed in advance rather than inferred from volume.

  • Stratify first, then size: quotas per cell before any prompt is written.
  • Fix the corpus for a measurement period; version it when it changes.
  • Replay each prompt multiple times per engine to separate signal from sampling noise.
  • Publish the frame alongside the score so the result can be audited.

Limitations

The figures in this paper are modelled and used to demonstrate the design; they are not client results. Engine behaviour changes between versions, so any absolute number ages faster than the method does.

Frequently asked

Is this peer reviewed?

No. It is a working paper published for scrutiny. We welcome replication attempts and will version the paper when the method changes.

Can I cite it?

Yes — cite as CiteRank AI Research (2026), including the version number shown on this page.

Methodology, limitations and disclosure

Every CiteRank study states who produced it, what it measured and where it stops being reliable. The full scoring model is documented on the methodology page.

Author
CiteRank AI Research
Author role
CiteRank AI Research team — measurement, prompt-corpus design and scoring
Review
Internal editorial review by the CiteRank AI Research team. No external or academic peer review was conducted.
Published
18 Jun 2026
Last updated
18 Jun 2026
AI engines
Not engine-specific
Model versions
Specific model build identifiers are not disclosed by every vendor and are therefore not claimed here.
Industry scope
Cross-industry
Geographic scope
Global
Language scope
English
Sample size
5 — Used to estimate run-to-run variance
Conflicts of interest
CiteRank AI publishes this research and sells an AI visibility platform. No third party funded, commissioned or reviewed this entry.
Data availability
Working paper (v1.2) — available on request.

Limitations

  • Figures in this entry are modelled and clearly labelled illustrative. They demonstrate structure and method; they are not observed client results.
  • AI engines are non-deterministic: an identical prompt can return a different answer on replay, so every figure is a sampled estimate rather than a fixed value.
  • Engine vendors change retrieval and ranking behaviour without notice. Findings describe the sampling window stated above, not a permanent state.
  • Results describe the prompt corpus that was designed for this study. A different corpus for the same brand can produce a materially different picture.

Suggested citation

CiteRank AI Research (2026). Sampling Bias in Prompt Corpora for Generative Engine Evaluation. CiteRank AI. https://www.citerank.in/research/paper-sampling-bias-prompt-corpora

Run Free Audit