The AI Search
Visibility Index
A study measuring which B2B brands AI assistants actually name and cite when buyers ask for recommendations. The protocol is published below in full. The findings are not, because we have not collected them yet.

Why this page has a method and no numbers
Publishing a study protocol before collecting data is what stops findings being reshaped to suit the publisher. This page documents the prompt set, engine coverage, repetition count, coding scheme and reporting plan for the AI Search Visibility Index in advance. Results will be added here when the first collection cycle completes — including any that are inconvenient for us.
Marketing research has a credibility problem, and most of it comes from findings published without a method anyone can inspect. A statistic with no stated sample, no coding rule and no repetition count is an assertion wearing a chart. We would rather publish the design and wait than fill the space with an estimate that gets cited as fact for the next three years.
The protocol, in full
Registered publicly so it cannot be adjusted after the fact. Any revision before collection will be dated and noted on this page.
| Design element | Specification | Reason |
|---|---|---|
| Unit of observation | One engine response to one prompt on one run | Rates require a defined denominator |
| Repetition | Five runs per prompt per engine | Generative output varies run to run |
| Engine coverage | ChatGPT Search, Perplexity, Gemini, AI Overviews, Copilot | The surfaces B2B buyers actually use |
| Interface | Consumer-facing, not API | API responses differ from what buyers see |
| Session state | Clean sessions, no personalisation carry-over | Personalisation contaminates comparison |
| Prompt set | Fixed across the study, frozen before collection | Comparability across brands and over time |
| Mention coding | Brand named anywhere in generated prose | Distinct from citation, moves independently |
| Citation coding | Linked attribution resolving to brand domain | Drives referral traffic; mention alone does not |
| Prominence coding | Ordinal position of first mention in answer | Buyers act on the first few names |
| Brand selection | By category, incumbents and challengers together | Avoids confirming that big brands are visible |
| Minimum reporting threshold | Stated with results; thin cells reported as thin | Prevents precision the sample cannot support |
This mirrors the protocol we run for individual clients, documented in full on the GEO testing methodology page. The index applies the same method across many brands rather than one.
What buyers actually type
The prompt set is built from how people describe problems, not from keyword volume. Keyword-derived prompts produce a study about SEO, not about AI recommendation.
Problem-led prompts dominate
Assistant queries run long and situational — a constraint, a budget, a stack. They rarely resemble the short category terms search keyword tools surface.
Frozen before collection
Once the set is fixed it does not change for the edition. Adjusting prompts mid-study is the most common way a visibility benchmark becomes meaningless.
Neutral phrasing
No prompt names a tracked brand unless the category makes it unavoidable. Brand-primed prompts inflate mention rates and tell you nothing about discovery.
Weighting is a design decision, published here so readers can judge whether it suits the questions they want answered.
What the first edition will contain
Committed in advance, so the eventual report cannot quietly drop a section that came out badly.
Mention rate distribution
How often tracked brands are named, reported as a distribution across the cohort rather than a single headline average.
Citation rate distribution
How often a linked attribution appears, reported separately from mentions because the two diverge substantially.
Prominence patterns
Where brands land in recommendation ordering, and whether position is stable across repeated runs.
Engine divergence
How far the five engines disagree with each other on the same prompt — potentially the most useful finding for practitioners.
Run-to-run variance
How much a single brand’s visibility moves across five identical runs, which sets the floor on what any measurement can claim.
Source composition
Which types of domain the engines actually retrieved from — owned sites, review platforms, press, forums.
What this study will not prove
Declared before collection, so they cannot be quietly omitted from the write-up.
It is a sample, not a census
A fixed prompt set cannot represent the unbounded range of wordings real buyers use. Findings are directional evidence about a defined set, not population parameters.
It cannot establish causation
Correlations between brand characteristics and visibility will not prove that changing one produces the other. Engines retrain on their own schedule and we do not control the environment.
It is a moment, not a trend
A first edition captures one window. A model update can reset the landscape between editions, and trend claims require several cycles run identically.
It is blind to personalisation
Clean-session results may differ from what a logged-in buyer with extensive chat history sees. Personalisation is growing and we cannot observe it.
Category coverage will be partial
The first edition will cover a limited set of B2B categories. Findings will not generalise to categories outside it, and the report will say which those are.
It measures presence, not persuasion
Being named is not being chosen. The study records visibility, and cannot tell you whether a mention influenced a purchase decision.
Two ways to be involved
Neither costs anything, and neither buys a favourable result — inclusion in the tracked set does not influence how a brand scores.
Request inclusion in the tracked set
Your brand is measured alongside category peers. You receive your own results ahead of publication, and decide whether they are attributed to you publicly or reported anonymously.
Get notified on publication
No drip sequence and no sales follow-up attached. One email when the first edition publishes, with the aggregate findings and the full dataset.
Participation terms
- Cost to participate
- None
- Influence on your result
- None
- Attribution
- Your choice
- Early access
- Yes
- Opt-out before publication
- Any time
- Sales follow-up
- Only if requested
About the index
What exactly will the AI Search Visibility Index measure?
Across a fixed set of commercial prompts and several AI engines, it records how often each tracked brand is named in the generated answer, how often it receives a linked citation, and where in the answer it appears. Mention and citation are recorded separately because they behave differently and have different downstream effects.
Why run each prompt multiple times?
Because generative engines are non-deterministic — the same prompt returns different brands on different runs. Single-pass measurement produces an anecdote rather than a rate. Repeated runs are what turn an observation into a frequency that can be compared over time.
How will brands be selected for the tracked set?
By category rather than by size alone, so each tracked category includes established incumbents alongside smaller challengers. Selecting only large brands would produce a study that confirms large brands are visible, which nobody needs.
Will you name brands with poor visibility?
Not without permission. Aggregate and category-level findings will be published openly; brand-level results will only be attributed where the brand has agreed. Publishing a named company’s weak result without consent is not research, it is a marketing stunt with a citation.
When will the first edition publish?
The protocol is live now and the first data collection cycle is being scheduled. We are deliberately not announcing a publication date we cannot yet commit to — a missed date on a research page undermines the credibility the research is meant to build.
How is this different from AI visibility tools already on the market?
Most commercial tools measure a single brand against a prompt set the customer defines, which is useful operationally but not comparable across companies. This index uses one fixed prompt set across many brands, which is what makes cross-brand and cross-category comparison meaningful.
Related material
Sources & references
- Google Search Central, “AI features and your website” — how AI Overviews source and link content.
- OpenAI, ChatGPT Search documentation — browsing behaviour and source attribution.
- Perplexity, publisher documentation and citation behaviour.
- Microsoft, Copilot search and attribution documentation.
- Oneskai GEO testing methodology (protocol GEO-M v4) — the per-client protocol this index scales.
- Study protocol registered on this page prior to data collection; revisions will be dated here.
Be measured, not guessed at
Request inclusion in the tracked set, or ask to be notified when the first edition publishes. If you would rather not wait, the same protocol runs as a standalone baseline for your brand today.