How we measure
AI visibility — in full
Asking ChatGPT about your brand once is not measurement. This page documents the protocol we use to produce a defensible number: which prompts, which engines, how many repetitions, what gets counted, and what we refuse to claim.
What GEO measurement actually is
GEO measurement is the repeated sampling of AI engine answers to a fixed prompt set, coded for whether a brand is mentioned, cited, and how prominently. Because generative outputs vary run to run, a single check is anecdote. Oneskai runs 100 prompts across 5 engines, 5 times each, producing 2,500 observations per cycle and a frequency rather than a claim.
The same prompt gives different answers
Generative engines sample from a probability distribution. Ask twice, get two answers. This is not a bug to work around — it is the property that makes single-shot checking worthless.
We have watched a brand appear in a recommendation list on one run and vanish on the next, minutes apart, same account, same wording. Any agency screenshotting a favourable ChatGPT response as proof of “AI visibility” is showing you one sample from a distribution they have not measured.
Five runs per prompt is the point where our observed brand-mention rate stabilises enough to be worth reporting, without the cycle cost becoming unreasonable. It is a practical trade-off, and we state it plainly rather than implying statistical certainty we do not have.
Read any single screenshot sceptically
Including ours. A result that cannot be reproduced across repeated runs should not appear in a report as evidence.
Observed pattern from a live engagement, brand anonymised. Reported mention rate for this prompt: 60%, not “we rank #2”.
Six steps, run the same way every cycle
Comparability is the whole point. If the method drifts between months, the trend line is meaningless.
Build the prompt set
Draft 100 prompts from sales call language, support tickets, category phrasing and objection questions. Review with your team. Freeze for the engagement.
Fix the environment
Clean sessions, no personalisation carry-over, logged location recorded, engine mode noted. Environment drift contaminates comparisons more than most people expect.
Execute 2,500 runs
100 prompts × 5 engines × 5 repetitions. Full response text captured verbatim alongside every source link the engine surfaced.
Code every observation
Each response coded for mention, citation, prominence position, competitor set and which of your URLs (if any) was retrieved.
Compute the seven metrics
Aggregate to rates, not counts. Report confidence qualitatively where the sample is thin rather than implying precision we cannot support.
Diff against baseline
Compare to the frozen kickoff cycle. Flag prompts that moved, prompts that decayed, and new competitors that entered the consideration set.
The seven metrics, defined precisely
Vague definitions are how visibility reports become unfalsifiable. Here is exactly what each number counts and what it does not.
| Metric | What it counts | What it does not count | Unit |
|---|---|---|---|
| Brand mention rate | Your brand named anywhere in the generated answer text | Links without the name appearing in prose | % of runs |
| Citation rate | A linked source attribution pointing at your domain | Being named without a link | % of runs |
| Citation prominence | Ordinal position of your first mention in the answer | Total number of mentions in one answer | Mean rank |
| Query coverage | Prompts where you appear in at least one of five runs | How often you appear within those prompts | % of prompt set |
| Source inclusion | Which specific URLs the engine retrieved from your domain | Domain-level mentions with no page attributed | URL list |
| AI referral traffic | Sessions arriving from AI engine referrers in analytics | Direct traffic from users who retyped your URL | Sessions |
| Pipeline attribution | CRM opportunities whose first touch was an AI referral | Assisted influence on deals sourced elsewhere | Count / value |
Mention rate and citation rate move independently. A brand can be recommended constantly and linked rarely — which produces authority without traffic, and needs a different fix.
Five engines, tested as buyers use them
We test the consumer interface rather than the API. API responses are cleaner and cheaper to collect, and they are not what your buyer sees.
ChatGPT Search & Perplexity
Browsing-enabled conversational engines. Highest citation density, most volatile between runs, and the two where source-link quality matters most.
Google AI Overviews & Gemini
Tied to Google’s index and quality systems, so they correlate more closely with classic organic authority than the standalone chat engines do.
Microsoft Copilot
Disproportionately relevant for enterprise and regulated buyers whose IT policy restricts them to Microsoft tooling during working hours.
Prompt set composition
100PROMPTSDefault weighting. Adjusted per engagement — category-creation businesses shift weight toward problem-led prompts because the category name does not exist yet.
What this method cannot tell you
Every measurement protocol has a boundary. Publishing ours is the difference between a methodology and a sales deck.
It is a sample, not a census
We observe 2,500 responses. Your buyers generate an unknown number in wordings we never tested. The number is directional evidence, not a population parameter.
It cannot prove causation
When mention rate rises after we ship work, that is correlation in an environment we do not control. Engines retrain and re-rank on their own schedule.
It cannot promise a placement
No input guarantees an output here. We improve the conditions that correlate with citation and report what happened, including when nothing did.
It is blind to personalisation
Engines increasingly personalise on account history. Our clean-session results may differ from what a logged-in buyer with long chat history sees.
It lags engine changes
A model update can reset the landscape between cycles. We flag suspected model shifts in the report rather than presenting the drop as a performance failure.
It does not measure sentiment depth
We code whether you were named and how prominently, not the full nuance of how you were characterised. Qualitative review is a separate, manual pass.
We do not optimise for an AI trick. We optimise the information ecosystem these systems read from — clear entities, extractable passages, original evidence, credible corroboration and clean technical access — and then we measure whether it moved.Oneskai GEO operating principle · reviewed quarterly
Protocol questions
Why test the same prompt five times?
Large language models are non-deterministic. The same prompt sent to the same engine minutes apart can return different brands, different orderings and different citations. A single check tells you what happened once, not what typically happens. Five runs per prompt lets us report a frequency rather than an anecdote.
How do you choose the 100 prompts?
We build the set from commercial reality, not keyword volume. Roughly half come from how buyers describe their problem, a quarter from category and comparison phrasing, and a quarter from objection and evaluation questions surfaced in sales calls. The set is frozen for the engagement so results stay comparable month to month.
Which engines do you test?
ChatGPT Search, Perplexity, Google Gemini, Google AI Overviews and Microsoft Copilot. We test the consumer-facing configuration rather than the API, because that is what buyers actually use. Where an engine offers both a browsing and non-browsing mode, we record which mode produced each result.
Can you make an AI engine cite our brand?
No. No agency can, and any guarantee of a citation should be treated as a red flag. Retrieval and generation are controlled by the engine and are probabilistic. What we can do is improve the inputs that correlate with citation — entity clarity, extractable passages, original evidence and third-party corroboration — and then measure honestly whether it moved.
How do you tell a mention apart from a citation?
A mention is your brand named in the generated text. A citation is a linked source attribution pointing at your domain. They move independently: brands are frequently named without being linked, and occasionally linked without being named in the prose. We report both separately because they have different downstream traffic effects.
What is citation prominence and why does it matter?
Prominence records where you appear in the answer: named first in a recommendation list, mentioned mid-list, or referenced only in a footnote. Buyers act disproportionately on the first two or three names in a generated recommendation, so moving from position six to position two matters more than adding a sixth mention elsewhere.
How long does a full benchmark cycle take?
A complete cycle is 100 prompts across 5 engines at 5 runs each — 2,500 observations — and takes about five working days to execute and code. We run a full cycle at kickoff to establish the baseline, then monthly thereafter, with the baseline frozen so later cycles remain comparable.
Do you share the raw results?
Yes. Clients receive the coded observation set, not just the summary chart, including the prompts where the brand did not appear at all. The prompts that return nothing are usually the most useful part of the dataset because they show where the content or corroboration gap actually is.
Related material
Sources & references
- Google Search Central, “AI features and your website” — how AI Overviews select and link sources.
- Google Search Central, “Creating helpful, reliable, people-first content” — E-E-A-T guidance.
- OpenAI, ChatGPT Search documentation — browsing behaviour and source attribution.
- Perplexity, publisher and citation documentation.
- Schema.org, Organization and Dataset vocabulary specifications.
- Oneskai internal protocol GEO-M v4, revised quarterly. Available to clients on request.
Get your baseline cycle
We build the 100-prompt set with your team, run all 2,500 observations, and hand back the coded dataset with your current mention rate, citation rate and prominence per engine — including every prompt where you did not appear.