Why an annual benchmark exists
Classical search analytics answer a question that is quietly becoming the wrong one: where does my page sit in a list of links? Generative engines do not return lists. They return an answer, and inside that answer a small number of brands are named, cited, or quietly omitted. The omission is invisible in every tool built for the link era.
This report exists to make that invisible layer measurable and repeatable. Every figure in it is derived from a designed prompt corpus, sampled on a fixed cadence, scored with a published model, and reported with its uncertainty. Where figures are modelled rather than observed from client workspaces, they are labelled illustrative.
The corpus: what was asked, and how it was built
A prompt corpus is not a keyword list with question marks bolted on. Real users state a situation, a constraint and a decision they are trying to make. The corpus is therefore constructed from intent archetypes rather than search volume.
- ›Discovery — 'who are the best X for Y' with no brand named.
- ›Shortlisting — 'compare A, B and C for a team of 30'.
- ›Validation — 'is A trustworthy / regulated / a good fit for us'.
- ›Displacement — 'alternatives to A' and 'A vs the market'.
- ›Operational — 'how do I do X', where a brand can be cited as the authority rather than the product.
Observed engine behaviour
The eight engines are not eight copies of one system. They differ in retrieval strategy, in how aggressively they cite, and in how much weight they place on freshness versus consensus.
- ›ChatGPT — strong prior-knowledge answers, citations appear mostly when browsing is triggered by recency or specificity.
- ›Google Gemini and AI Overviews — heavily grounded in the live index; entity data and structured markup pay off fastest here.
- ›Anthropic Claude — conservative naming, favours brands with corroborated, non-promotional sources.
- ›Perplexity — the most citation-dense engine; being a cited source matters more than being the recommended brand.
- ›Microsoft Copilot — index-grounded with an enterprise skew; documentation and compliance pages carry unusual weight.
- ›xAI Grok — the most recency-sensitive; social and news signals move answers within days.
- ›Meta AI — conversational and brevity-biased; typically names two to three brands maximum.
The scoring model in one page
Every sampled answer is parsed into three observable events: the brand is ranked (given an explicit position), cited (a link or source attribution), or mentioned (named without attribution). Those events are weighted, normalised by prompt-set size, and rolled into the AI Visibility Score.
Recommendation Share answers a narrower and more commercially useful question: of all the brands named in answers to your prompt set, what proportion of the naming is you? It is share-of-voice for the answer layer, and it is the metric most tightly coupled to pipeline in AI-native categories.
What actually moved the numbers
Across the programmes observed, three interventions produced repeatable movement, and several popular ones produced none.
- ›Entity hygiene — one canonical name, address, description and identifier set, corroborated across independent sources.
- ›Extractable structure — claims stated once, plainly, near a heading that matches the question being asked.
- ›Third-party corroboration — engines weight what others say about you above what you say about yourself.
- ›Interventions with minimal observed effect: keyword density, page count growth, and AI-generated content published at volume.
How to read this report without over-fitting
Benchmarks describe a population, not your brand. Two companies in the same category can sit twenty points apart on Visibility Score for reasons that have nothing to do with content quality — a merged entity record, a rebrand that never propagated, or a category term the engines interpret differently.
Use the benchmark to size the gap. Use your own prompt corpus to explain it.
