First principle: measure the answer, not the page
Every design decision follows from one observation — the unit of competition is the answer, and inside an answer a brand is either present or absent. Position in a link list is not a proxy for that.
Prompt-set design
A prompt corpus is built from the buyer's decision path, then stratified so no single intent dominates the score. Corpora are frozen for the duration of a measurement window; changing prompts mid-window makes trend lines meaningless.
- ›Stratify across discovery, shortlisting, validation, displacement and operational intents.
- ›Include unbranded prompts — a corpus of branded prompts measures recall, not competitiveness.
- ›Localise where the buying decision is local; language and region change the answer.
- ›Version the corpus, and report the version alongside the score.
Sampling under non-determinism
Ask the same engine the same question twice and you may get two different brand lists. This is not noise to be eliminated; it is the property being measured. Each prompt is therefore sampled repeatedly across a window, and results are reported as frequencies.
A single screenshot proves nothing. A frequency across repeated samples is evidence.
Parsing: what counts as what
Each response is parsed into structured events before scoring.
- ›Ranked — the brand is given an explicit ordinal position or presented as the primary recommendation.
- ›Cited — the answer attributes a source to the brand's domain or a document it authored.
- ›Mentioned — the brand is named without attribution or ordering.
- ›Absent — the brand does not appear, which is itself a recorded observation.
Scoring and normalisation
Events are weighted (ranked above cited above mentioned), normalised by the size and composition of the prompt set, and combined into the AI Visibility Score on a 0–100 scale. Engine scores are reported separately as well as combined, because an engine-level move and a market-level move demand different responses.
What the score does not claim
It is not a traffic forecast, it is not a causal model, and it is not comparable across differently designed prompt corpora. Two brands are only comparable when measured on the same corpus, in the same window, in the same locale.
