Define the unit first
AI visibility is measurable only after a team defines the unit of observation. Here, one unit is a saved prompt run with engine, date, relevant conditions, answer text, visible citations, and a coding rule. The following metrics describe that sample. They do not expose an engine’s internal ranking and they should never be generalized from one prompt run to an entire market. Freeze definitions before reporting so a score change is not merely a change in coding.
Answer inclusion rate
Formula: prompt runs that name the brand or include its page divided by eligible prompt runs, multiplied by 100. Interpretation: frequency of appearance in this defined sample. Failure mode: combining prompts where inclusion is impossible with prompts where it is natural, or changing the rule mid-period. Hypothetical example: 8 of 20 eligible runs include the brand, so inclusion rate is 40%. That describes those 20 observed runs only; it says nothing about all users, all prompts, or future answers.
Citation share
Formula: visible citations to your domain divided by all visible citations in eligible answers, multiplied by 100. Interpretation: share of shown source links in the sampled outputs. Failure mode: treating a link as endorsement without checking whether the target page supports the nearby claim. Hypothetical example: 6 of 60 visible citations point to your domain, producing 10%. Also record page-level targets, because a domain may be cited for material irrelevant to the decision you care about.
Share of voice
Formula: your coded brand mentions divided by total coded mentions across a defined competitor set, multiplied by 100. Interpretation: relative presence when answers name one of the selected brands. Failure mode: using an incomplete peer set or treating repeated mentions in one response as independent demand. Hypothetical example: your brand receives 9 mentions and four peers receive 36 combined, for 20% share of voice. Preserve the peer list, matching rules, and prompt set for every reporting period.
Average mention position
Formula: sum of the brand’s ordinal positions in eligible explicitly ordered lists divided by its eligible mentions. Interpretation: placement where the interface actually presents a sequence. Failure mode: inventing rank from prose that is not a list, or ignoring ties and grouped recommendations. Hypothetical example: positions 2, 1, and 4 average 2.33. Only use this measure when the visible answer creates an order; a paragraph mentioning three options does not reliably supply one.
Sentiment distribution
Formula: for each of positive, neutral, negative, and mixed, divide the count of that coded category by all actual coded mentions, then multiply by 100. Report no-mention rate separately: eligible prompt runs with no brand mention divided by eligible prompt runs, multiplied by 100. Interpretation: the sentiment distribution describes language about the brand when it appears; no-mention rate describes absence from the defined prompt sample. Failure mode: adding absent runs to the sentiment denominator, or substituting reviewer intuition for a written codebook. Hypothetical example: of 10 actual mentions, 2 positive, 6 neutral, 1 mixed, and 1 negative yield 20%, 60%, 10%, and 10%. If 10 of 20 eligible runs have no brand mention, the separately reported no-mention rate is 50%; it is not a sentiment category. Keep the excerpt and a rationale for each label so disagreement remains visible.
Source-domain concentration
Formula: citations from the most-cited domain divided by all visible citations, multiplied by 100; optionally report the top three domains. Interpretation: how concentrated the observed source set is. Failure mode: calling concentration quality or bias without reading the sources; a narrow factual query may reasonably draw on one authoritative domain. Hypothetical example: 30 of 50 citations come from one domain, so top-domain concentration is 60%. Inspect page relevance before proposing a content or technical action.
AI-referred sessions or conversions
Formula: sessions or conversions with an attributable AI referrer divided by all sessions or conversions, multiplied by 100, using consented analytics and a documented attribution rule. Interpretation: observed referral activity, not the value of every answer mention. Failure mode: assuming every product sends a stable referrer or treating missing referral data as zero influence. Hypothetical example: 30 attributable AI-referred sessions from 3,000 total sessions equals 1%; two later conversions should be reported with the attribution window, not as causal proof.
Use the metrics responsibly
Freeze a core prompt set, record engine and conditions, retain raw outputs where policy permits, and publish coding definitions next to each chart. Google and OpenAI document access foundations, while citation research shows why visible links need close reading. These measures help notice a change and choose an investigation. They cannot guarantee visibility, diagnose a hidden algorithm, or turn one observation into a benchmark. Start with the Small-Team AI Visibility Audit so the underlying records remain inspectable.
Set denominator rules before collecting results
Most misleading visibility charts fail in the denominator. Before a run, identify which prompts are eligible for each metric and why. A question asking for a definition may be eligible for source-domain concentration but not for brand inclusion. A response that declines to answer may be retained as an observed outcome but excluded from a citation-share denominator under a written rule. Do not remove inconvenient outputs after seeing them. If a rule changes, restate the prior period where feasible or show a clear break in the series.
Keep the sampling frame beside the number: prompt text, intent class, engine or surface, language, geography if known, date, account state if known, and the coding version. Aggregate only comparable observations. A 40% inclusion rate from twenty product-comparison prompts is not directly comparable with 40% from twenty broad educational prompts, even when the arithmetic matches. Likewise, an engine that displays many source cards can produce a different citation-share denominator from one that shows one or none. The metric is a description of a defined measurement procedure, not a portable property of the brand.
For practical reporting, pair every percentage with its raw count and an uncertainty note. “8 of 20 eligible saved runs, core English sample, July observation” is harder to overinterpret than “40% AI visibility.” It also tells a teammate how to repeat the work. This does not require a large research program; it requires resisting the urge to hide a small or changing denominator behind a polished headline.
Read metrics together, not as a league table
Each measure answers a different question. Inclusion rate may rise because a brand is mentioned in a caveat. Citation share may fall because an interface shows more sources overall. Share of voice depends on the selected peer set. Average position is meaningful only for genuinely ordered lists. Sentiment depends on a transparent codebook. Concentration can be appropriate for a narrow factual topic. Referred sessions depend on consent, attribution, and whether a product passes a stable referrer. A dashboard should therefore put the underlying excerpts and raw counts within reach of the chart.
Hypothetical example: a team sees inclusion rise from 8 of 20 to 11 of 20 saved runs while citation share stays at 6 of 60 visible links and referred sessions remain flat. The defensible interpretation is limited: the brand appeared more often in these comparable saved runs. It would be wrong to announce a traffic or revenue lift, or to conclude that citations caused the mentions. The next sensible check is qualitative: read the new mentions, verify whether they are favorable or qualified, and confirm that the sample and coding rules did not change.
Use metrics to select a question for investigation, then return to the pages and evidence. Google’s documentation makes clear that ordinary crawlability and textual accessibility remain relevant to its AI Search features, while the underlying source-selection logic is not reduced to a public scorecard. That is why a metric should lead to an inspectable action—repair a page, clarify a claim, update a source, or improve a measurement rule—rather than a promise that a particular result will recur.
For a weekly scorecard, show the sample size, the full eligibility rule, the observed count, the formula, and a link to the saved evidence. Keep a separate field for changes in prompts, interfaces, tracking, or coding. A chart that makes its assumptions visible is slower to produce, but it is far more useful when someone asks whether an apparent change is real.
Related reading: The Small-Team AI Visibility Audit.
Methodology
Metrics are proposed operational definitions for repeated manual observations. They are not platform-reported metrics and do not establish causal ranking effects.
Sources
- Google Search CentralAI Features and Your Website(opens in a new tab)
- Google Search CentralRobots Meta Tag Specifications(opens in a new tab)
- OpenAI Help CenterAdvertiser Guidance for Allowing OpenAI Web Crawlers(opens in a new tab)
- Perplexity DocumentationPerplexity Crawlers(opens in a new tab)
- arXivRetrieval-Augmented Generation for Knowledge-Intensive NLP Tasks(opens in a new tab)