Skip to content
CiteRank AI
Dataset

GEO Prompt Corpus — Sample Release

A documented sample of the stratified prompt corpus used in CiteRank benchmarks, with intent class, buying stage and geography labelled on every row.

CiteRank AI Research·2 July 2026·8 min readIllustrative data
DatasetPromptsOpen Data

Key takeaways

  • Every row carries its stratification labels, so you can reproduce our quotas or build your own frame from them.
  • No client data appears in any release — rows are modelled from public category language.
  • The schema is versioned; breaking changes get a new major version rather than a silent overwrite.
Resource

Sample dataset

Placeholder dataset containing modelled rows. Access is granted on request while the schema is stabilised.

Rows
2,500 (sample of 40k)
Format
CSV + JSONL
Licence
CC BY 4.0
Updated
Quarterly

Schema

One prompt per row, with the labels needed to rebuild the sampling frame.

  • prompt_id — stable identifier across releases.
  • prompt_text — the natural-language prompt as issued to the engine.
  • intent_class — informational, comparative, advisory or transactional.
  • buying_stage — awareness, evaluation or decision.
  • geo — country or region the prompt is scoped to.
  • category — industry vertical.

Intended use

Build a comparable corpus for your own category, audit our published benchmarks, or teach prompt-corpus design without needing engine access.

Methodology, limitations and disclosure

Every CiteRank study states who produced it, what it measured and where it stops being reliable. The full scoring model is documented on the methodology page.

Author
CiteRank AI Research
Author role
CiteRank AI Research team — measurement, prompt-corpus design and scoring
Review
Internal editorial review by the CiteRank AI Research team. No external or academic peer review was conducted.
Published
2 Jul 2026
Last updated
2 Jul 2026
AI engines
Not engine-specific
Model versions
Specific model build identifiers are not disclosed by every vendor and are therefore not claimed here.
Industry scope
Cross-industry
Geographic scope
Global
Language scope
English
Sample size
Not disclosed for this entry
Conflicts of interest
CiteRank AI publishes this research and sells an AI visibility platform. No third party funded, commissioned or reviewed this entry.
Data availability
Sample dataset — available on request.

Limitations

  • Figures in this entry are modelled and clearly labelled illustrative. They demonstrate structure and method; they are not observed client results.
  • AI engines are non-deterministic: an identical prompt can return a different answer on replay, so every figure is a sampled estimate rather than a fixed value.
  • Engine vendors change retrieval and ranking behaviour without notice. Findings describe the sampling window stated above, not a permanent state.
  • Results describe the prompt corpus that was designed for this study. A different corpus for the same brand can produce a materially different picture.
  • This entry does not disclose a sample size, so its figures should not be treated as statistically representative.

Suggested citation

CiteRank AI Research (2026). GEO Prompt Corpus — Sample Release. CiteRank AI. https://www.citerank.in/research/dataset-geo-prompt-corpus-sample

Run Free Audit