Every enterprise now claims a “GEO strategy”. Almost none can produce a baseline. When a board asks how much AI visibility has changed since an engagement began, the honest answer at most agencies is a screenshot of a single ChatGPT response, taken on a single day, with no record of what the same prompt returned a month earlier. That is not measurement. It is anecdote dressed as evidence, and it collapses the moment a CFO asks for the methodology behind it.
Deloitte’s guidance on AI initiative measurement is blunt on this point: the absence of a locked, pre-deployment baseline is the single largest source of indefensible ROI claims. GEO is no exception. If you cannot show what citation share looked like before an engagement, you cannot show what changed because of it. This article sets out the measurement framework Daryo89 uses to establish that baseline, define what actually counts as a citation, and report citation velocity in a way that survives scrutiny from a technically literate audience rather than collapsing under it.
Why a Single-Engine Snapshot Is Not a Baseline
AI citation is not stable. Research into citation churn indicates that 40–60% of citations for a given query set can rotate month to month, driven by model updates, re-crawling cycles, and shifts in which sources the retrieval layer deems current or authoritative. A brand cited in a ChatGPT response on a Tuesday may be absent by the following Tuesday, not because anything about the brand’s content changed, but because the retrieval index reshuffled. Any framework that treats a single query, on a single engine, on a single day, as a meaningful data point is measuring noise and calling it signal.
The correction is structural, not clever. A defensible baseline requires a fixed library of prompts, run across multiple engines, on a repeating cadence, with every response logged and parsed before any engagement work begins. This is the pre-deployment lock Deloitte describes applied to a GEO context: the baseline is captured, frozen, and treated as the reference point against which all subsequent measurement is compared. Without that lock, “improvement” is simply whatever number happens to be more favourable on the day someone chooses to report it.
Constructing the Prompt Library
The prompts used for baseline measurement must be brand-neutral. This is a deliberate and non-negotiable constraint: querying “tell me about [Brand]” tests whether an AI model has indexed a brand’s existence, not whether it recommends the brand organically to a genuine buyer. The prompts that matter are the ones a prospective customer would actually type — “best commercial construction firms in Cambridgeshire for a design-and-build contract” rather than “what does [Company] do”. Branded prompts inflate visibility scores artificially and produce a baseline that flatters the client rather than describing the market.
A working prompt library for a B2B professional services client typically spans several clusters: category-defining queries (the sector and service type without any brand name), comparison queries (structured as “X vs Y” or “best options for Z”), problem-first queries (framed around the buyer’s pain point rather than the solution category), and due-diligence queries (the kind a procurement function or agentic buying assistant would run before shortlisting a vendor). Each cluster behaves differently across engines, and tracking them separately — rather than as one blended average — is what allows the eventual velocity reporting to mean anything.
Multi-Engine Querying
The prompt library is then run identically across ChatGPT, Claude, Gemini and Perplexity, and, where relevant, Google’s AI Mode surface. This is not a stylistic choice — it reflects the fact that these systems draw on materially different retrieval architectures, different underlying indexes, and different citation display conventions, and they do not converge on the same sources. Independent analysis has found that only around 11% of domains cited by ChatGPT are also cited by Perplexity for comparable queries. A brand that has engineered strong visibility in one engine and assumes that generalises to the others is working from a false premise. The baseline has to be per-engine, not blended, or the framework will conceal exactly the divergence it exists to reveal.
Benchmark citation rates also differ meaningfully by platform, which matters when interpreting a baseline in isolation. Reported 2026 benchmarks put typical citation rates at roughly 2.0+ per relevant query set for ChatGPT, 1.5–2.0 for Perplexity, and a comparatively lower 1.1–1.5 for Google’s AI Mode. A client scoring 1.3 on Google AI Mode is not necessarily underperforming — it may simply be operating within the expected range for that surface. Reporting raw counts without engine-specific context invites exactly the kind of overstated or understated claim this framework is designed to prevent.
Citation Versus Mention: A Distinction That Determines the Entire Report
Before any frequency, velocity or share-of-voice figure can be calculated, the terms have to be fixed precisely, because the two most commonly conflated outcomes in AI-generated answers behave in entirely different ways.
A citation is an evidence-level attribution: the AI system explicitly credits a specific, retrievable webpage as the source for a claim, typically rendered as an inline numbered reference, a follow-up source list, an embedded hyperlink, or a visual source card. Perplexity and Google’s AI Overviews favour inline numbered citations; ChatGPT more commonly surfaces a follow-up source list. In every case, the mechanism is the same — the model has actively retrieved a specific document during response generation and is pointing the user back to it. This is a verifiable, checkable claim of provenance.
A mention is narrative-level visibility without attribution: the brand or entity name appears in the answer text, but no link, source card or explicit “according to” phrasing accompanies it. Mentions frequently originate not from live retrieval but from the model’s training data — residual signal from social discussion, forum threads, press coverage or historic indexing that shaped the model’s parametric knowledge long before the query was asked. A mention tells you the model has heard of you. A citation tells you the model trusted your page enough, at query time, to point a user directly at it.
This distinction is not academic. Data on generative search outputs indicates that brand mentions occur roughly 3.2 times more frequently than citations — mentions are the much larger, noisier signal, and citations are the rarer, harder-won one. But the commercial value runs in the opposite direction: AI search visitors arriving via a genuine citation convert at rates reported to be as much as 4.4 times higher than typical organic search visitors, and AI-referred traffic more broadly converts in the 14–40% range depending on sector and intent. A report that tallies “AI visibility” as mentions and citations combined is averaging together a high-volume, low-value signal with a low-volume, high-value one, and will systematically misrepresent commercial impact in either direction depending on which happens to dominate that month.
Any GEO reporting framework Daryo89 produces logs these as two separate line items, never one blended score. A client whose mention count rises while citation count stagnates has gained ambient familiarity but not retrieval trust — a meaningfully different, and less commercially useful, outcome.
Time-to-Citation as a Distinct Metric
Frequency answers “how often are we cited”. It does not answer “how quickly does new or refreshed content earn a citation once published”. That second question is what we term time-to-citation: the interval, measured in tracking cycles rather than calendar guesswork, between a piece of content going live (or being materially restructured) and its first appearance as a cited source in the tracked prompt set.
This metric matters because it is one of the few genuinely leading indicators available in GEO measurement. Citation frequency at any given moment reflects the cumulative effect of everything published to date; time-to-citation isolates the responsiveness of the retrieval layer to a specific, dated change. A shorter time-to-citation after a structural intervention — semantic chunking, explicit entity schema, answer-first content restructuring — is meaningful evidence that the intervention affected retrieval behaviour, in a way that an aggregate frequency count, muddied by unrelated content already in the index, cannot isolate.
Time-to-citation is also where domain authority becomes visible as a confound rather than a variable to ignore. Content published by lower-authority or newer domains is documented to take measurably longer to surface in citations than equivalent content from established, high-trust domains, independent of the content’s structural quality. This means time-to-citation figures should never be reported as an absolute benchmark (“content gets cited within N days”) but always as a relative, within-domain comparison: this domain’s time-to-citation before the engagement, versus this domain’s time-to-citation after it, holding the domain constant. Comparing a client’s time-to-citation against an industry-wide average invites exactly the kind of confounded, indefensible claim the framework exists to prevent.
Perplexity’s retrieval behaviour offers a useful mechanical illustration of why structure interacts with speed. Perplexity is understood to visit approximately ten candidate pages per query but cite only three to four of them. Time-to-citation, in this context, is partly a function of whether a page’s answer-first framing — content structured so the direct answer sits within the first meaningful third of the page — allows it to clear the retrieval-to-citation threshold on an early visit, rather than being crawled repeatedly without ever being selected as one of the cited few.
Calculating Share of Voice Across Engines
Share of voice in a GEO context is the proportion of citations, within a defined prompt cluster and engine, attributed to a given domain relative to all domains cited across that same cluster and engine over the same tracking window. It is deliberately not a proportion of all possible answers — most answers in a well-constructed neutral prompt set will cite no single brand exclusively, and several will cite none of the tracked competitor set at all. Share of voice is calculated only across the subset of responses that produced at least one citation, which keeps the metric meaningful rather than diluted by prompts where citation behaviour was never in play.
Share of Voice (engine, cluster) = citations attributed to Domain A ÷ total citations across all tracked domains, within that engine and prompt cluster, over the tracking window.
The calculation must be run separately per engine and per prompt cluster, then reported as a matrix rather than a single blended figure. A domain with strong share of voice in the “problem-first” cluster on Perplexity but negligible presence in the “comparison” cluster on ChatGPT is telling you something operationally specific: the content answering direct problem queries is structurally sound and retrievable, but comparison-format content — the kind procurement functions and B2A agents actually query during due diligence — is either absent or not structured for citation. That is an actionable finding. A single blended share-of-voice percentage across all engines and clusters conceals it entirely.
Because of the 40–60% monthly citation churn already noted, share of voice should be reported as a rolling trend across a minimum of several tracking cycles, not as a point-in-time figure. A single week’s share-of-voice reading is close to meaningless in isolation — it can move sharply purely from index churn unrelated to any content change. The trend line across weeks is where volatility smooths into a signal that can be reported with confidence.
Reporting Velocity Without Overstating Causality
This is the section where most GEO reporting fails, because it is where the temptation to claim a direct causal line — “we published this page, therefore the AI now recommends you” — is strongest, and the underlying evidence rarely supports that strength of claim. Large language models identify statistical correlation and retrieval-relevance patterns; they do not perform causal inference, and neither, in most cases, does the measurement infrastructure sitting on top of them. Causal AI as a distinct discipline exists precisely because standard LLM-based systems do not do this natively, and it remains a nascent, data-hungry field with its own interpretive risks, including the reinforcement of existing biases in the training corpus.
A defensible framework therefore separates metrics into leading and lagging indicators, and reports each with the appropriate level of confidence rather than treating them interchangeably. Leading indicators — citation visibility rate by engine and country, the specific mix of URLs being cited, performance by prompt cluster, and the count of newly gained versus newly lost citations week over week — tell you whether an intervention is producing a measurable effect on retrieval behaviour before any downstream commercial effect has had time to surface. Lagging indicators — referral traffic attributable to AI engines, and conversions assisted by AI-sourced visits — confirm business impact, but they surface later and are more exposed to confounding from unrelated marketing activity.
The methodologically honest position is to report leading indicators as evidence of association, using language such as “citation share increased in the weeks following the restructuring of X content, consistent with improved retrieval of that page” rather than “this content caused the citation increase”. Where causality claims are genuinely warranted — for instance, where a controlled before/after comparison isolates a single structural change with no other confounding activity in the same window — that isolation should be stated explicitly, including what was held constant and what was not. Comparative baseline methodologies of this kind, including neutral-statement testing under varied attribution conditions to quantify bias in AI outputs, as used in University of Zurich research in 2025, offer a template for the level of rigour this requires: define the control condition, define the variable being tested, and report only what the design actually supports.
The Reporting Discipline That Makes This Defensible
Put together, the framework produces a report structure that looks nothing like a marketing dashboard: a locked pre-engagement baseline across all four engines and every prompt cluster; citation and mention counts tracked as separate, never-blended line items; time-to-citation measured within-domain against that domain’s own pre-engagement figure, never against an external benchmark; share-of-voice calculated per engine and per cluster and read as a rolling trend rather than a snapshot; and velocity reported through explicitly labelled leading and lagging indicators, with causal language reserved only for designs that can actually support it.
The value of this discipline is not that it produces more impressive numbers. Often it produces less impressive ones — a flat trend line, a citation share that moved within the range explainable by ordinary monthly churn, a time-to-citation improvement too small to distinguish from noise. That is the point. A framework that can report a null result honestly is the only kind of framework whose positive results are worth believing. For technically literate boards, the credibility of the measurement methodology is now inseparable from the credibility of the GEO strategy itself.
Frequently asked questions
Why is a single AI query snapshot not a valid GEO baseline?
Research into citation churn indicates that 40-60% of citations for a given query set can rotate month to month due to model updates and re-crawling, so a single query on a single day measures noise rather than a stable signal.
What is the difference between a citation and a mention in AI-generated answers?
A citation is an explicit, verifiable attribution where the AI system links to a specific retrieved webpage, while a mention is the brand name appearing in the answer text without a link, often surfaced from training data rather than live retrieval.
Why must GEO be measured across multiple AI engines rather than just one?
Independent analysis found that only around 11% of domains cited by ChatGPT are also cited by Perplexity for comparable queries, meaning strong visibility in one engine does not generalise to others and baselines must be tracked per-engine.
What is time-to-citation and why does it matter?
Time-to-citation is the interval between content going live or being restructured and its first appearance as a cited source, and it functions as a leading indicator because it isolates the effect of a specific change rather than blending it into an aggregate frequency count.