AI Search measurement

How we measure AI Search without turning it into a vanity score.

AI Search measurement is a repeated sample of defined buyer questions across named engines, markets and dates. We keep brand mentions, competitor share, answer position, citations, retrieval, sentiment, factual accuracy and commercial outcomes separate. A change in one metric does not stand in for another, and movement after our work does not by itself prove causation.

Method updated 12 September 2026

The dolphin harbourmaster checks five answer systems and their sources

The short version

What does a credible AI Search report need to show?

A credible report shows the question set, engines, country, language, sample dates and valid-answer count. It defines every metric before showing a score. It also marks missing runs and method changes, because a neat percentage built on changing coverage is not a clean comparison.

We use scheduled observations to catch changes, then compare completed periods defined in the written measurement plan. The buyer can see what moved, what evidence supports the interpretation, what remains unavailable and which next action is justified.

01

What exactly do we sample?

The measurement frame is agreed before the baseline. Each response remains an observation tied to its settings, not a timeless ranking.

Measurement tools and source maps arranged on a shelf

Buyer questions

A versioned set of real category, comparison, problem and buying questions. Branded checks are reported separately so they do not inflate discovery visibility.

Engines and surfaces

Only the named engines and answer surfaces in scope. Results are broken out by engine before any combined view.

Market and language

Country, language and locale are fixed where the engine allows it. A UK English result is not silently mixed with a German result.

Time and coverage

Before the baseline, the plan states the exact repeat count for each question, engine, market and language cell, the schedule, the period length and the minimum valid coverage needed for comparison. Every run records its date and outcome.

Evidence

We retain the answer, brand and competitor mentions, visible order, linked or cited sources, and the factual statements needed for accuracy review.

Calculation rules

A valid answer is a completed in-scope response that can be stored and reviewed. The plan freezes cell weights, inclusion rules, metric formulas, evidence authority, low-volume threshold and any uncertainty method before the baseline.

02

What do the seven measurement layers mean?

Each layer answers a different question. The report keeps the numerator, denominator and unavailable observations visible so the score can be checked.

LayerDefinitionWhat it can tell youWhat it cannot prove
VisibilityValid answers naming the brand at least once divided by all valid answers in the frozen cell set. One answer contributes at most one brand-presence count.Whether the brand appears for the tracked question set.Prominence, preference, a citation or a visit.
Share of voiceBrand-presence credits divided by all brand and frozen-competitor presence credits. Each tracked brand contributes at most one credit per answer; the report shows both counts.Relative presence inside the defined competitor set.Market share, buyer preference or revenue share.
PositionFor valid answers that both name the brand and contain a meaningful ordered list, record the first visible brand position. Report the eligible-answer count, position distribution and median, not a rank for ineligible answers.Whether the brand tends to appear early or late when order is observable.A universal rank. Many answers have no meaningful ordered list.
Citations and retrievalA citation is a visible supporting link. Retrieval is a separately recorded source-use event only when the measurement system exposes it, whether or not that source is also cited. Retrieved-but-not-cited is a reported subset.Which domains and pages support or are exposed as inputs to sampled answers, and where source gaps exist.That source use caused the brand mention, or that an unexposed retrieval event occurred.
SentimentEach brand-bearing answer is labelled positive, neutral, negative or mixed under the versioned rubric. The report shows category counts and shares, retains the passage and sends ambiguous cases for review.Whether recurring themes appear in the sample.Customer satisfaction or reputation across the whole market.
AccuracyCorrect checkable claim occurrences divided by all checkable claim occurrences about the brand. The plan names the approved source of truth; disputed and uncheckable claims stay outside the denominator.Where engines repeat outdated, incomplete or incorrect information.That silence is accurate, or that a correct claim persuaded the buyer.
Qualified outcomesA funnel, not one blended score. Report separate counts and conversion rates for identifiable engaged visits, enquiries, booked calls and opportunities that pass the pre-agreed qualification rule, with unknown source retained as unknown.Whether measurable commercial activity followed the exposure path.That AI visibility alone caused the outcome.

03

How do we compare movement without hiding drift?

AI answers can change when models, source indexes, location, personalisation or the question wording change. We therefore preserve the measurement frame and report the actual coverage beside every comparison.

A planning table with versioned questions and comparison periods
  1. 1. Freeze the frame

    Version the questions, competitors, engines, markets, cell weights, repeats, period length, formulas and decision threshold before the baseline.

  2. 2. Record every outcome

    Count valid answers and label unavailable, failed or unsupported observations. Missing data is not zero.

  3. 3. Aggregate repeated samples

    Give each frozen cell its predeclared weight. Use 95% Wilson intervals only for unweighted answer-level proportions that meet the plan's independence rule. Correlated repeats or weighted estimates need a separately frozen cluster-aware or design-based variance method; otherwise label the interval unavailable and the estimate descriptive.

  4. 4. Segment before combining

    Inspect engine, topic, funnel stage and market separately. A combined score can hide one strong engine and one failing engine.

  5. 5. Annotate interventions

    Record when a page, source, prompt set, tracking rule or model changes. The timeline makes interpretation possible without pretending it isolates cause.

04

What is the difference between observation, attribution and causation?

The strongest claim must match the strongest evidence available.

Observed

A metric moved after a dated change. This is a signal worth investigating.

Attributed

A visit or outcome carries a source, campaign or declared discovery path under the chosen analytics rules. Credit depends on the attribution model and lookback window.

Corroborated

Answer visibility, referred visits and qualified outcomes move in a consistent direction across repeated periods, with major confounders reviewed.

Causal

The evidence isolates the intervention from other plausible causes, usually through a suitable experiment or control. A before-and-after chart alone does not meet this bar.

05

How do AI Search metrics connect to qualified commercial outcomes?

We join the journey only where the evidence allows it. Answer data shows exposure. Analytics shows identifiable visits and on-site actions. Forms and booking systems show enquiries and appointments. Qualification belongs in the CRM or an agreed review, not in a pageview count.

A signal station connects answers, visits and qualified outcomes
01

Exposure

Brand visibility, share of voice, position and supporting sources.

02

Engagement

Identifiable referred visits, methodology-page engagement and audit transitions.

03

Intent

Submitted audits, contact forms or booked-call starts with safe campaign fields.

04

Qualification

Accepted fit criteria, opportunity status and disqualification reason where available.

05

Outcome

Won work or another agreed commercial result, reported separately from attributed influence.

06

What happens when data is unavailable?

We say unavailable and explain the boundary. We do not turn a failed connector, unsupported engine field, blocked report or missing booking classification into zero. The denominator changes only when the report shows the change.

  • Search Console includes traffic from Google's AI features in the Web search type, so it does not provide a standalone AI-feature traffic total. Conversion attribution belongs in analytics and the CRM.
  • ChatGPT search referral links can carry utm_source=chatgpt.com, but citations without a click never appear as referral sessions.
  • An engine may show a linked source without exposing whether or how it was retrieved internally.
  • Attribution settings and lookback windows change how analytics assigns credit. The report names the rule used.
  • Low volume can make a percentage unstable. We show counts and avoid a commercial conclusion until the evidence is large enough for the decision.

07

Which public sources define the reporting boundaries?

These sources explain platform mechanics. They do not endorse Schmitdy or guarantee visibility.

Marco Lobo
See the questions and evidence behind your AI Search baseline.The free audit maps your buyer questions, current answer coverage and the clearest source gaps. No invented baseline and no promise that one edit caused the result.
Get your free AI Search audit β†’

FAQ

Questions a sceptical buyer should ask

Is AI Search visibility the same as website traffic?+

No. Visibility measures whether a brand appears in sampled answers. Traffic measures identifiable visits. An answer can influence a buyer without a click, while a visit can arrive without a measured brand mention.

Can share of voice be compared across two different prompt sets?+

Not cleanly. Share of voice depends on the questions, engines, competitor set, market and valid-answer coverage. If any of those change, the report must label the new version and avoid presenting the movement as a like-for-like trend.

Does an early position mean the engine recommends the brand?+

Not necessarily. Position records visible order only when order is meaningful. Recommendation strength, wording, citations and accuracy need separate review.

Is a citation the same as retrieval?+

No. A citation is a visible supporting link. Retrieval is an internal use of a source and is recorded only when the measurement source exposes it. We never infer hidden retrieval from a mention alone.

How is sentiment checked?+

The report retains the relevant answer passage, assigns a defined label and reviews recurring themes by topic. The score is an interpretation of the sample, not a customer-satisfaction survey.

What counts as an accurate answer?+

A checkable statement must match current approved evidence. Outdated, incomplete and incorrect statements are separated, and disputed or uncheckable claims stay outside the accuracy denominator.

Can a rise in AI visibility be credited with new revenue?+

Only to the level the evidence supports. Tagged visits, declared discovery paths and CRM qualification can support attribution. A visibility increase followed by revenue is still not causal proof without a design that rules out other plausible causes.

What should appear in the first report?+

The first report should show the frozen measurement frame, baseline counts, metric definitions, missing-data rules, source examples, qualification rules and the next observation window. If the current baseline is unavailable, it should say unavailable rather than record zero.

Your existing stack

Keep the platform. Make the measurement inspectable.

The method works with the website, analytics and CRM stack you already use. We agree the evidence path before claiming an outcome.

SquarespaceVercelWebflowSanityShopifyHubSpot

Next step

See the questions and evidence behind your AI Search baseline.

The free audit maps your buyer questions, current answer coverage and the clearest source gaps. No invented baseline and no promise that one edit caused the result.

Get your free AI Search audit β†’