Methodology

Every AI answer starts with a prompt.

Scroll through the story of how a single question becomes a reproducible visibility score across the 8 engines shaping how buyers discover you.

Research vs. Demo Lab

CiteRank maintains a strict separation between Original Research (verified population-level studies with disclosed samples) and the Demo Lab (modelled scenarios and synthetic dashboards). Research methodology described here applies to all verified audits and population benchmarks.

Methodology Version: 2.2 · Last updated: August 12, 2026 · View Changelog

Authors: Shivaji Alaparthi, Anudeep Manchineni Published: August 12, 2026 Modified: August 12, 2026
01Chapter 1 — The prompt graph

Buyers don't ask one question.

Buyers don't ask one question — they traverse a funnel. We build a category-specific prompt graph from your industry, sub-vertical, geography and competitor set. Every prompt is tagged by buying-stage intent — so the scoring weights match what actually drives pipeline. CiteRank currently uses a proprietary 35/30/20/15 intent weighting. This is a measurement assumption designed to reflect different stages of buyer intent and is periodically reviewed as search behaviour changes. Our model reflects the broader market trend toward conversational and AI-assisted discovery (Gartner 2024: <a href='https://www.gartner.com/en/newsroom/press-releases/2024-02-19-gartner-predicts-search-engine-volume-will-drop-25-percent-by-2026-due-to-ai-chatbots-and-other-virtual-agents' target='_blank' rel='noopener noreferrer' class='underline hover:text-teal-400'>Search Engine Volume Predict</a>).

Discovery
"best AI SEO tool for B2B SaaS"
intent weight: 35%Broad category research and awareness.
Comparison
"best AI SEO platform for performance teams"
intent weight: 30%Feature-by-feature side-by-side evaluation.
Evaluation
"does citerank support Gemini and Claude"
intent weight: 20%Entity-specific trust and capability check.
Conversion
"citerank pricing for marketing teams"
intent weight: 15%High-intent transactional information gathering.
02Chapter 2 — Multi-LLM execution

The same prompt, replayed 10–30 times across 8 engines.

LLMs are non-deterministic — ask the same question twice and you'll get different answers. We replay every prompt 10–30 times per engine to control for temperature and answer drift, then average. Per plan: Platform 10 · Platform + GEO 20 · Managed GEO 30; Enterprise scales further on request. 95% confidence intervals are reported where the measurement design supports them.

8 engines × 12 replays shown
ChatGPTGeminiClaudePerplexityGoogle AIOCopilotGrokMeta AI
cited in response not cited
03Chapter 3 — Why one query isn't enough

Answer drift is real. Confidence bands make it honest.

We report every score with a 95% confidence interval where the measurement design supports them so single-run noise doesn't drive decisions. If the band is wide, you know before you act.

Visibility per replay
Same prompt · Same engine · 10 runs
52 ± 9
95% CI
run 1run 10
04Chapter 4 — Citation detection

Every URL, every mention, resolved to one entity.

Responses are parsed for explicit citations, inline brand mentions and entity-level references. Aliases, typos and domain variants all collapse to the same brand so nothing is double-counted or missed.

parsed response
citerank.in
URL citation
CiteRank AI
brand mention
CiteRank
alias resolved
cite-rank
typo variant
competitor.io
entity separated

One brand, many surfaces.

We deduplicate near-mentions and resolve corporate aliases so "CiteRank", "CiteRank AI" and the .in domain all match a single entity. Sentiment and recommendation-order are extracted per mention.

05Chapter 5 — The composite score

Four signals become one score, 0–100.

Intuition first, math second. Each signal has a plain-language explanation — expand any card to see the formula, why it exists and how it's weighted.

share_b = Σ mentions_b / Σ mentions_all
Weighting:45%

Why it exists — Baseline signal — presence in the answer at all.

order_b = Σ (1 / position_b,i)
Weighting:33%

Why it exists — Buyers act on what's recommended first.

sent_b ∈ [-1, +1]
Weighting:22%

Why it exists — Being cited badly is not the same as being cited well.

Composite
score_b = Σ_p w_p · ( α·share_b,p + β·order_b,p + γ·sent_b,p )

The Reconciled Composite (Methodology 2.2): To ensure the scoring matches pipeline value, we use normalized weights of α=0.45, β=0.33, and γ=0.22 (summing to 1.0). Unlike "black-box" alternatives, every component is traceable and reproducible.

The Multiplicative Volatility Penalty (w_p)

Rather than an additive 10% weight, CiteRank uses a multiplicative per-prompt reliability multiplier. This down-weights high-variance (noisy) prompts where LLM temperature leads to inconsistent answers.

Stable Prompt (σ=0.1)
w_p = 1/(1+0.1) = 0.91
Retains 91% of measured signal
Noisy Prompt (σ=0.6)
w_p = 1/(1+0.6) = 0.62
Heavily penalized (38% reduction)

*Ref: Uncertainty-aware measurement standards in Sielinski et al. (arXiv 2603.08924).

06Chapter 6 — Gaps, re-runs, attribution

A score without a fix is a vanity metric.

Every brand score has matched competitor scores. Where you lose, we ship a gap-fix brief — the specific prompts you don't appear on, the sources cited instead, and the entity, schema or content changes that would close the gap.

Gap-fix briefs

Prompt-level briefs listing the sources currently cited and the schema/entity moves to close the gap.

Standard Re-runs

Automated monthly full-graph runs for Platform. Platform + GEO and above include weekly 150-prompt partial runs (pulses) so you see movement between major reports.

Downstream attribution

Optional GA4, Search Console and HubSpot links relate AI visibility shifts to sessions and pipeline. Observed relationships are not proof of causation.

07Chapter 7 — Who it's for

Built for teams who need to defend a number.

Marketing leaders

Measure AI visibility movement for the board with reproducible scores and CIs.

SEO / GEO teams

See exactly which prompts and sources are winning — and what to ship next.

Agencies

White-label reports, per-client prompt graphs, and gap-fix briefs baked in.

Product & PR

Track how launches and press cycles land in the actual AI answer surface.

What we don't do

No single-shot prompts. No black-box scoring. No fabricated samples.

  • · Every score is averaged across 10–30 replays per engine (Platform 10 · Platform + GEO 20 · Managed GEO 30).
  • · The composite formula and weights are published in-product.
  • · Marketing samples are labelled "Illustrative"; customer reports use only that customer's real runs.
Run Free Audit