GEO benchmarks 2026: AI search visibility.
A reproducible way for B2B teams to measure whether their brand is retrieved, represented accurately, cited, and connected to commercial outcomes across AI answer systems.
What changed in AI search after June 2026.
The interfaces changed, but the publisher task remains stable: allow the relevant crawler, keep the page indexable, publish useful evidence, and measure each answer surface separately. No platform documents a special HTML format for GPT-5.6, Claude Opus 5, Claude Fable 5, or Google AI Overviews.
| Answer surface | Documented discovery layer | What a publisher should do | What is not supported |
|---|---|---|---|
| ChatGPT search, including GPT-5.6 sessions | OAI-SearchBot and OpenAI crawler controls | Permit search crawling, keep important evidence available without a login, and track ChatGPT referral traffic separately | A GPT-5.6-only page variant or special ranking schema |
| Claude search and agent use, including Opus 5 and Fable 5 sessions | Claude-SearchBot and Claude-User controls | Permit the search and user agents needed for the intended use; keep claims easy to verify from direct sources | Separate copy for each Claude model family |
| Google AI Overviews and AI Mode | Google Search index, retrieval-augmented generation, and query fan-out | Meet normal Search technical requirements, remain snippet eligible, and publish original, non-commodity value | Special GEO markup, mandatory AI chunking, or llms.txt as a Google Search signal |
Google added generative AI performance reporting in June 2026 and began testing publisher controls for inclusion in Search generative features. Record those platform changes in the study log so a reporting change is not mistaken for a content effect.
Measure the full path from retrieval to revenue.
| Stage | Metric | What it answers |
|---|---|---|
| Retrieval | Prompt coverage | For what share of the fixed prompt panel does the brand or its evidence appear? |
| Representation | Answer accuracy | Are the brand, offer, location, proof, and limitations described correctly? |
| Selection | Mention and citation share | How often is the brand named, and how often is a supporting source linked? |
| Source influence | Cited-source distribution | Which owned and third-party pages shape the answer? |
| Commercial action | AI referrals and assisted conversions | Did a user visit, return, start a form, or become a qualified opportunity? |
| Business outcome | Pipeline and revenue | Which opportunities and revenue can be directly or assistively attributed? |
A seven-step repeatable baseline.
- Define the commercial decision. State the audience, category, geography, and business action the prompt panel represents.
- Build a versioned prompt panel. Start with 30 to 100 recommendation, comparison, alternatives, pricing, implementation, problem, and risk questions derived from real sales and search evidence.
- Fix the test conditions. Record engine, model or interface where known, account state, geography, language, device, and collection date.
- Repeat each prompt. Run at least five observations per engine and period. A single stochastic answer is not a trend.
- Capture the raw answer and sources. Record every named brand, order, cited URL, source type, accuracy issue, and whether a clickable referral is present.
- Score only after retaining the observations. Aggregate by prompt and intent cluster; never discard the rows behind a composite score.
- Connect to commercial systems. Preserve AI referral UTMs, original landing pages, assisted journeys, qualified opportunities, proposals, and revenue.
Fields included in the CSV.
| Field group | Required fields | Purpose |
|---|---|---|
| Test identity | study_id, prompt_id, prompt_version, intent_cluster | Prevents prompt drift from being mistaken for performance change |
| Conditions | engine, model_or_surface, run_date_utc, geography, language, account_state, run_number | Makes observations reproducible and comparable |
| Answer | answer_text, brands_mentioned, target_brand_mentioned, target_brand_position | Records inclusion and recommendation context |
| Citations | cited_urls, target_domain_cited, citation_position, source_types | Separates brand knowledge from visible source attribution |
| Quality | accuracy_status, inaccurate_claims, analyst_notes | Measures whether visibility is helpful or harmful |
| Commercial | landing_sessions, qualified_leads, opportunities, attributed_revenue_usd | Connects visibility to business outcomes when data is available |
Download visibility-template.csv. The file contains headers, a field-definition row, and one clearly marked example row. It contains no client observations.
Keep the core metrics auditable.
- Prompt coverage
- Unique prompts with at least one target-brand mention divided by total prompts in the fixed panel.
- Mention rate
- Runs that mention the target brand divided by all valid runs.
- Citation rate
- Runs that cite the target domain divided by all valid runs. Report this separately from mention rate.
- Mention share of voice
- Target-brand mentions divided by all named-brand mentions in the same prompt panel and period.
- Citation share of voice
- Target-domain citations divided by all provider-domain citations in the same panel and period.
- Answer accuracy
- Valid target-brand mentions without a material factual error divided by all target-brand mentions.
A composite score is optional. If one is used, publish its weights and always report the underlying metrics. Changing weights changes the story without changing the observations.
Controls that stop false improvement claims.
- Freeze the core prompt panel before an intervention and version every addition or wording change.
- Use repeated runs and report both totals and unique-prompt coverage.
- Store the raw answer, cited URLs, date, engine, and run conditions.
- Separate branded prompts from non-branded recommendation and comparison prompts.
- Record model and interface changes; do not compare periods silently when the test surface changed.
- Label owned, partner, customer, independent editorial, review, community, and official sources separately.
- Do not refresh publication dates unless the content or underlying observations materially changed.
- Do not describe owned-company results as independent client validation.
What this framework can and cannot prove.
| It can support | It cannot establish by itself |
|---|---|
| A reproducible record of prompts, outputs, mentions, citations, sources, and accuracy | A universal ranking factor or stable cross-market benchmark |
| Before-and-after visibility comparison under stated conditions | That one page or schema property caused a model response |
| Referral and assisted-journey measurement | Revenue causality without opportunity-level attribution and controls |
| Source gaps and entity contradictions requiring action | Guaranteed recommendation or citation placement |
What current evidence supports and what it does not.
| Supported practice | Unsupported shortcut |
|---|---|
| Keep pages crawlable, indexable, snippet eligible, and accessible to the documented search crawler. | Treat llms.txt as a Google ranking factor or replace the normal Search index with it. |
| Publish original evidence, clear source links, direct answers, and claim-level limitations. | Change a date without a material review, create fake third-party mentions, or add unsupported superlatives. |
| Cover the real buyer decision with distinct intent, sector context, and implementation detail. | Generate many near-duplicate fan-out pages for lexical variants of the same intent. |
| Repeat a stable prompt panel and keep raw observations because engines, models, and runs vary. | Call one answer a benchmark, stable ranking, or causal result. |
| Use valid structured data that matches visible content. | Assume special schema or model-specific HTML guarantees a citation. |
Recent research also supports this cautious approach. Document-level properties can matter more than small lexical edits; decision-demand coverage can affect retrieval; generative systems vary across engines and time; and authority-aware retrieval matters most in high-stakes domains. See the ACL study on document properties, research on decision-demand coverage, cross-engine stability findings, and authority-aware retrieval research.
Use first-party platform data where it exists.
- Bing Webmaster Tools AI Performance reports citations, cited pages, grounding queries, and trends.
- OpenAI’s crawler documentation explains OAI-SearchBot, GPTBot, and ChatGPT-User controls.
- Anthropic’s crawler documentation explains Claude-SearchBot, Claude-User, and ClaudeBot controls.
- Microsoft’s AI-search content guidance recommends clear titles and headings, direct answers, lists, tables, and evidence.
- Google’s AI-search optimization guide says normal Search requirements apply and no special GEO schema or AI-specific content format is required.
AI-search measurement questions.
What is AI search visibility?
AI search visibility describes whether a brand appears accurately in answers generated from a defined set of buyer questions. It should be separated into mention share, citation share, source influence, answer accuracy, and referral behavior. A single visibility score can summarize trends, but the underlying prompt-level observations must remain available for audit.
How should AI search visibility be measured?
Use a fixed, versioned prompt panel; record engine, model, date, geography, run number, brand mentions, cited URLs, answer accuracy, and competitors; then repeat each prompt at least five times. Report results by intent cluster and retain the raw observations. Connect referral sessions and qualified opportunities separately because a citation is not a lead.
What is the difference between a mention and a citation?
A mention occurs when an answer names the brand. A citation occurs when the answer links or attributes a source supporting its response. The two must be measured separately: a brand can be mentioned from model knowledge without a visible source, while a page can be cited without the brand receiving a prominent recommendation.
Does schema markup guarantee an AI citation?
No. Structured data can clarify entities and page meaning, but it does not create authority or guarantee selection. Search and AI systems also evaluate crawlability, relevance, extractable evidence, freshness, and corroboration. Schema should mirror visible content exactly and be treated as one retrieval aid inside a broader technical and editorial system.
How many prompts belong in an AI visibility baseline?
Start with 30 to 100 commercially meaningful prompts grouped by recommendation, comparison, alternatives, pricing, implementation, risk, and problem intent. The right number depends on the market and available analyst capacity. Coverage matters more than volume: each prompt should represent a real buyer decision and have a documented inclusion reason.
How often should an AI visibility panel be repeated?
Run a stable panel monthly for strategic reporting and more frequently around a dated intervention. Keep the prompt wording, geography, account state, and run protocol fixed so periods remain comparable. Record model and date changes explicitly. Volatile outputs require repeated runs; a single answer is an observation, not a trend.
Can this framework prove that GEO caused revenue growth?
Not by itself. The framework can show changes in retrieval, mentions, citations, sources, accuracy, and referral behavior. A causal revenue claim requires a dated intervention, a prior baseline, stable measurement, opportunity attribution, and controls for other campaigns or market changes. Report assisted influence separately from directly attributable revenue.
Need an independent baseline and source strategy?
The commercial GEO/AEO service applies this framework to a company’s real buyer questions, technical retrieval, entity evidence, source influence, and pipeline attribution.
Review the GEO/AEO engagement scope