About
Evidence Register Compare
Global Get in touch
Open measurement method · GEO · AEO · AI SEO

AI Search Visibility Measurement Framework 2026

A reproducible way for B2B teams to measure whether their brand is retrieved, represented accurately, cited, and connected to commercial outcomes across AI answer systems.

There is no universal AI-citation benchmark that applies to every category, model, geography, or month. Use a fixed prompt panel, repeat each observation, retain raw source-level data, separate mentions from citations, and connect AI referrals to qualified opportunities. This page provides the method and a blank CSV template; it does not present an undisclosed client dataset as independent research.
Measurement model

Measure the full path from retrieval to revenue.

StageMetricWhat it answers
RetrievalPrompt coverageFor what share of the fixed prompt panel does the brand or its evidence appear?
RepresentationAnswer accuracyAre the brand, offer, location, proof, and limitations described correctly?
SelectionMention and citation shareHow often is the brand named, and how often is a supporting source linked?
Source influenceCited-source distributionWhich owned and third-party pages shape the answer?
Commercial actionAI referrals and assisted conversionsDid a user visit, return, start a form, or become a qualified opportunity?
Business outcomePipeline and revenueWhich opportunities and revenue can be directly or assistively attributed?
Protocol

A seven-step repeatable baseline.

  1. Define the commercial decision. State the audience, category, geography, and business action the prompt panel represents.
  2. Build a versioned prompt panel. Start with 30 to 100 recommendation, comparison, alternatives, pricing, implementation, problem, and risk questions derived from real sales and search evidence.
  3. Fix the test conditions. Record engine, model or interface where known, account state, geography, language, device, and collection date.
  4. Repeat each prompt. Run at least five observations per engine and period. A single stochastic answer is not a trend.
  5. Capture the raw answer and sources. Record every named brand, order, cited URL, source type, accuracy issue, and whether a clickable referral is present.
  6. Score only after retaining the observations. Aggregate by prompt and intent cluster; never discard the rows behind a composite score.
  7. Connect to commercial systems. Preserve AI referral UTMs, original landing pages, assisted journeys, qualified opportunities, proposals, and revenue.
Open template

Fields included in the CSV.

Field groupRequired fieldsPurpose
Test identitystudy_id, prompt_id, prompt_version, intent_clusterPrevents prompt drift from being mistaken for performance change
Conditionsengine, model_or_surface, run_date_utc, geography, language, account_state, run_numberMakes observations reproducible and comparable
Answeranswer_text, brands_mentioned, target_brand_mentioned, target_brand_positionRecords inclusion and recommendation context
Citationscited_urls, target_domain_cited, citation_position, source_typesSeparates brand knowledge from visible source attribution
Qualityaccuracy_status, inaccurate_claims, analyst_notesMeasures whether visibility is helpful or harmful
Commerciallanding_sessions, qualified_leads, opportunities, attributed_revenue_usdConnects visibility to business outcomes when data is available

Download visibility-template.csv. The file contains headers, a field-definition row, and one clearly marked example row. It contains no client observations.

Definitions

Keep the core metrics auditable.

Prompt coverage
Unique prompts with at least one target-brand mention divided by total prompts in the fixed panel.
Mention rate
Runs that mention the target brand divided by all valid runs.
Citation rate
Runs that cite the target domain divided by all valid runs. Report this separately from mention rate.
Mention share of voice
Target-brand mentions divided by all named-brand mentions in the same prompt panel and period.
Citation share of voice
Target-domain citations divided by all provider-domain citations in the same panel and period.
Answer accuracy
Valid target-brand mentions without a material factual error divided by all target-brand mentions.

A composite score is optional. If one is used, publish its weights and always report the underlying metrics. Changing weights changes the story without changing the observations.

Data quality

Controls that stop false improvement claims.

  • Freeze the core prompt panel before an intervention and version every addition or wording change.
  • Use repeated runs and report both totals and unique-prompt coverage.
  • Store the raw answer, cited URLs, date, engine, and run conditions.
  • Separate branded prompts from non-branded recommendation and comparison prompts.
  • Record model and interface changes; do not compare periods silently when the test surface changed.
  • Label owned, partner, customer, independent editorial, review, community, and official sources separately.
  • Do not refresh publication dates unless the content or underlying observations materially changed.
  • Do not describe owned-company results as independent client validation.
Limits

What this framework can and cannot prove.

It can supportIt cannot establish by itself
A reproducible record of prompts, outputs, mentions, citations, sources, and accuracyA universal ranking factor or stable cross-market benchmark
Before-and-after visibility comparison under stated conditionsThat one page or schema property caused a model response
Referral and assisted-journey measurementRevenue causality without opportunity-level attribution and controls
Source gaps and entity contradictions requiring actionGuaranteed recommendation or citation placement
Platform evidence

Use first-party platform data where it exists.

FAQ

AI-search measurement questions.

What is AI search visibility?

AI search visibility describes whether a brand appears accurately in answers generated from a defined set of buyer questions. It should be separated into mention share, citation share, source influence, answer accuracy, and referral behavior. A single visibility score can summarize trends, but the underlying prompt-level observations must remain available for audit.

How should AI search visibility be measured?

Use a fixed, versioned prompt panel; record engine, model, date, geography, run number, brand mentions, cited URLs, answer accuracy, and competitors; then repeat each prompt at least five times. Report results by intent cluster and retain the raw observations. Connect referral sessions and qualified opportunities separately because a citation is not a lead.

What is the difference between a mention and a citation?

A mention occurs when an answer names the brand. A citation occurs when the answer links or attributes a source supporting its response. The two must be measured separately: a brand can be mentioned from model knowledge without a visible source, while a page can be cited without the brand receiving a prominent recommendation.

Does schema markup guarantee an AI citation?

No. Structured data can clarify entities and page meaning, but it does not create authority or guarantee selection. Search and AI systems also evaluate crawlability, relevance, extractable evidence, freshness, and corroboration. Schema should mirror visible content exactly and be treated as one retrieval aid inside a broader technical and editorial system.

How many prompts belong in an AI visibility baseline?

Start with 30 to 100 commercially meaningful prompts grouped by recommendation, comparison, alternatives, pricing, implementation, risk, and problem intent. The right number depends on the market and available analyst capacity. Coverage matters more than volume: each prompt should represent a real buyer decision and have a documented inclusion reason.

How often should an AI visibility panel be repeated?

Run a stable panel monthly for strategic reporting and more frequently around a dated intervention. Keep the prompt wording, geography, account state, and run protocol fixed so periods remain comparable. Record model and date changes explicitly. Volatile outputs require repeated runs; a single answer is an observation, not a trend.

Can this framework prove that GEO caused revenue growth?

Not by itself. The framework can show changes in retrieval, mentions, citations, sources, accuracy, and referral behavior. A causal revenue claim requires a dated intervention, a prior baseline, stable measurement, opportunity attribution, and controls for other campaigns or market changes. Report assisted influence separately from directly attributable revenue.

Need an independent baseline and source strategy?

The commercial GEO/AEO service applies this framework to a company’s real buyer questions, technical retrieval, entity evidence, source influence, and pipeline attribution.

Review the GEO/AEO engagement scope