Problem statement
Generative engines answer questions, not queries. A corpus built from keyword exports therefore samples a different population than the one being measured, and the resulting score drifts whenever the engine reweights its retrieval.
This paper formalises that gap and proposes a stratified alternative that can be reproduced by a third party from the published design alone.
Proposed sampling design
Prompts are drawn from a frame defined by intent class, buying stage, geography and specificity, with quotas fixed in advance rather than inferred from volume.
- ›Stratify first, then size: quotas per cell before any prompt is written.
- ›Fix the corpus for a measurement period; version it when it changes.
- ›Replay each prompt multiple times per engine to separate signal from sampling noise.
- ›Publish the frame alongside the score so the result can be audited.
Limitations
The figures in this paper are modelled and used to demonstrate the design; they are not client results. Engine behaviour changes between versions, so any absolute number ages faster than the method does.
