Methodology

How We Measure AI Visibility

The prompt-set design, sampling cadence, parsing rules and scoring model behind the AI Visibility Score — including what the score deliberately does not claim.

CiteRank AI Research·5 January 2026·14 min readMethod note

Method note. Method note. This entry documents how CiteRank measures. It reports no findings; any numbers shown are worked examples.

MethodologyScoringSampling
Observable events
3
Ranked · Cited · Mentioned
Engines
8
Sampled independently, never averaged blind
Minimum samples
Repeated
Per prompt, per engine, across a window

Key takeaways

  • Engine outputs are non-deterministic, so every measurement is a distribution, not a value.
  • Three observable events — ranked, cited, mentioned — carry different weights and different meaning.
  • A score without its prompt corpus is uninterpretable; the corpus is part of the measurement.

First principle: measure the answer, not the page

Every design decision follows from one observation — the unit of competition is the answer, and inside an answer a brand is either present or absent. Position in a link list is not a proxy for that.

Prompt-set design

A prompt corpus is built from the buyer's decision path, then stratified so no single intent dominates the score. Corpora are frozen for the duration of a measurement window; changing prompts mid-window makes trend lines meaningless.

  • Stratify across discovery, shortlisting, validation, displacement and operational intents.
  • Include unbranded prompts — a corpus of branded prompts measures recall, not competitiveness.
  • Localise where the buying decision is local; language and region change the answer.
  • Version the corpus, and report the version alongside the score.

Sampling under non-determinism

Ask the same engine the same question twice and you may get two different brand lists. This is not noise to be eliminated; it is the property being measured. Each prompt is therefore sampled repeatedly across a window, and results are reported as frequencies.

A single screenshot proves nothing. A frequency across repeated samples is evidence.

Parsing: what counts as what

Each response is parsed into structured events before scoring.

  • Ranked — the brand is given an explicit ordinal position or presented as the primary recommendation.
  • Cited — the answer attributes a source to the brand's domain or a document it authored.
  • Mentioned — the brand is named without attribution or ordering.
  • Absent — the brand does not appear, which is itself a recorded observation.

Scoring and normalisation

Events are weighted (ranked above cited above mentioned), normalised by the size and composition of the prompt set, and combined into the AI Visibility Score on a 0–100 scale. Engine scores are reported separately as well as combined, because an engine-level move and a market-level move demand different responses.

What the score does not claim

It is not a traffic forecast, it is not a causal model, and it is not comparable across differently designed prompt corpora. Two brands are only comparable when measured on the same corpus, in the same window, in the same locale.

Frequently asked

Why not just average all eight engines?

Because the engines have different retrieval behaviour and different audiences. A blind average hides the engine-specific movement that is usually the actionable part.

How large should a prompt corpus be?

Large enough that no single intent stratum dominates and small enough to re-sample on cadence. Composition matters far more than raw count.

Evidence basis

How this entry was produced

Method note. This entry documents how CiteRank measures. It reports no findings; any numbers shown are worked examples.

Research type
Method note — documentation of measurement approach
Classification
Method note
Data collected
Published 2026-01-05. No data collection: this entry documents method rather than reporting a study.
Prompt sample size
Not applicable — no corpus was sampled for this entry.
Replays per prompt
Not applicable — no prompts were replayed for this entry.
Engines covered
Method applies to all eight supported engines: ChatGPT, Gemini, Claude, Perplexity, Google AI Overviews, Microsoft Copilot, Grok, Meta AI.
Engine versions
Not stated. Engine vendors do not expose a stable build identifier for every model, so a version cannot be claimed accurately.
Methodology
Written by the CiteRank research team from the production measurement pipeline. Internal editorial review only; no external or academic peer review.
Limitations
  • This entry reports no findings. Any number shown is a worked example chosen for clarity, not a measurement.
  • AI engines are non-deterministic: an identical prompt can return a different answer on replay, so any figure derived from them is an estimate, never a fixed value.
  • Engine vendors change retrieval and ranking behaviour without notice, so anything stated here describes the stated window only.

Methodology, limitations and disclosure

Every CiteRank study states who produced it, what it measured and where it stops being reliable. The full scoring model is documented on the methodology page.

Author
CiteRank AI Research
Author role
CiteRank AI Research team — measurement, prompt-corpus design and scoring
Review
Internal editorial review by the CiteRank AI Research team. No external or academic peer review was conducted.
Published
5 Jan 2026
Last updated
5 Jan 2026
AI engines
Not engine-specific
Model versions
Specific model build identifiers are not disclosed by every vendor and are therefore not claimed here.
Industry scope
Real Estate
Geographic scope
Global
Language scope
English
Sample size
Repeated — Per prompt, per engine, across a window
Conflicts of interest
CiteRank AI publishes this research and sells an AI visibility platform. No third party funded, commissioned or reviewed this entry.
Data availability
Underlying raw data is not published. Method and scoring are documented on the Methodology page.

Limitations

  • AI engines are non-deterministic: an identical prompt can return a different answer on replay, so every figure is a sampled estimate rather than a fixed value.
  • Engine vendors change retrieval and ranking behaviour without notice. Findings describe the sampling window stated above, not a permanent state.
  • Results describe the prompt corpus that was designed for this study. A different corpus for the same brand can produce a materially different picture.

Corrections and revisions

No corrections have been issued for this entry since publication on 5 Jan 2026. If a figure or claim here is wrong, write to research@citerank.in. Substantive corrections are published inline with the date they were made, and the original wording is retained in the note.

Suggested citation

CiteRank AI Research (2026). How We Measure AI Visibility. CiteRank AI. https://www.citerank.in/research/how-we-measure-ai-visibility

Run Free Audit