The two-hour question

A small-team audit should answer a modest question: what did a defined set of answer experiences show today, and which accessible evidence gap is worth fixing? It should not claim to measure the whole AI market, predict revenue, or prove a ranking mechanism. Two focused hours can produce a useful baseline when prompts, conditions, outputs, and limitations are written down. The value is comparability over time, not a perfect score from a single session.

Build a prompt sample

Start with 12 to 20 prompts. Cover problem exploration, category comparisons with named criteria, alternatives, implementation questions, and brand questions. Tag each by funnel stage, geography, language, audience, and intent. Use customer language from interviews, sales notes, support, and search queries where permission allows. This is an operationally useful sample, not a statistically representative sample of every user. Record why each prompt entered the set so future reviewers can see the selection method.

Manual engine checks

Allocate 45 minutes. For every prompt, record engine, date and time, signed-in or signed-out state where known, geography or language setting, exact wording, full answer, visible citations, and whether the response declined or asked a follow-up. Do not keep retrying until a preferred answer appears; that creates selection bias. If personalization or interface variation cannot be controlled, record the limitation beside the result. Preserve raw output only where your privacy and platform rules permit it.

Capture brand mention

Code whether the organization is named and in what role: recommended option, example, comparison point, warning, or absent. A name without a source can affect a buyer’s shortlist, but it should not be counted as a citation. Keep a short quotation or faithful excerpt to justify the code. Define a brand and product-name matching rule in advance, including how to treat parent companies, abbreviations, and ambiguous terms. That makes the next audit comparable rather than dependent on memory.

Capture citation, sentiment, competitors

For citations, save the target page and decide whether it supports the nearby claim. For sentiment, label wording positive, neutral, negative, mixed, or absent with a short rationale; this is coding, not mind reading. For competitors, record every named alternative and the criterion used to describe it. Do not transform a few mentions into market share. The records should show what the answer displayed, the conditions of the observation, and the uncertainty around its interpretation.

Review access foundations

Spend 30 minutes on pages you hoped a reader could find. Confirm a public successful response, no unintended noindex, appropriate robots access, canonical consistency, descriptive internal links, and important information as text. Google recommends these fundamentals for AI features and says structured data must match visible text. OpenAI documents separate risks from robots rules, WAFs, CAPTCHAs, login, and rate limits. Passing a check reduces an avoidable issue; it does not establish inclusion in an answer engine.

Use a priority matrix

Score every finding by user impact, evidence confidence, effort, and reversibility. High impact, high confidence, low effort goes first: for example, restoring a blocked documentation page or adding an omitted limitation to an authoritative guide. High impact but low confidence becomes an investigation; perhaps a source is absent but the sample is too small to explain why. Low impact and high effort belongs in a backlog. The matrix prioritizes decisions and never claims that a page edit will cause a future citation.

Weekly cadence

Monday: run a stable core of five prompts. Tuesday: check cited pages against original sources. Wednesday: fix one access or evidence gap. Thursday: review consented analytics for attributable referred sessions or conversions, without inferring sessions from citations. Friday: summarize what changed and what remains uncertain. Once a month, rerun the broader sample under comparable conditions. Hypothetical example: an answer names a company but cites a general guide; its own page lacks formulas and a date. Improving that page is a high-confidence content action, not a prediction of inclusion.

A two-hour run of show

Prepare before the timer starts. Put the prompt list, matching rules, company and competitor aliases, a spreadsheet or note template, and the priority URLs in one place. Write the session date, intended locale, language, and account state at the top. This prevents the audit from spending its first half hour deciding what a “mention” means. If a prompt is sensitive, uses customer information, or could expose confidential plans, remove or rewrite it; an observation routine is not a reason to disclose private data to a third-party service.

Minutes 0–15: confirm the sample and the coding rules. Minutes 15–60: run the selected prompts across the chosen visible answer experiences, saving outputs consistently and recording unavailable or declined responses rather than omitting them. Minutes 60–80: code brand appearance, cited URLs, competitor names, ordered-list position when genuinely present, and the sentiment wording. Minutes 80–105: inspect the priority pages for the access and evidence basics already listed. Minutes 105–120: assign each finding an owner, due date, confidence note, and next observation date.

The timeboxes are a guardrail, not a claim of scientific precision. A new team can start with fewer prompts or one carefully documented surface. The important constraint is comparability: run the same core prompts under similar recorded conditions next time, and keep expansions to the sample labeled as expansions. A prompt sample selected from actual buyer questions may be commercially useful, yet it still does not represent all possible users or all outputs from an engine.

Make the coding reviewable

A codebook turns a collection of screenshots into a reusable record. Define “brand included” before looking at the results: for example, exact product name, approved aliases, and how parent-company references are handled. Define a citation as a visible link or source card to a particular URL, not merely a company name in prose. Define sentiment from the language in the response—such as an explicit praise, limitation, warning, or neutral description—and permit “mixed” and “unclear” rather than forcing a positive or negative label.

Add a short evidence field for every code. For a mention, preserve the surrounding sentence. For a citation, save the target URL and the claim it appears to support. For a competitor, save the criterion named beside it. A second reviewer need not agree with every judgment, but should be able to see how it was made. If two reviewers disagree, record the disagreement and revise the rule if necessary; silently averaging incompatible interpretations makes a future trend line less trustworthy.

The same discipline applies to technical findings. A page that returns a successful response may still contain a noindex directive, a restrictive preview control, a confusing canonical, or essential content that is unavailable in initial text. Record the actual condition and the test used, not a generic “SEO issue.” Google’s guidance distinguishes eligibility from any guarantee of crawl, index, or serving, so the audit should label an access finding as an avoidable barrier rather than evidence of the hidden reason for a missing answer.

Turn the matrix into a small decision log

Use a four-cell action note for every material finding: observation, evidence, next action, and uncertainty. Example: observation—three saved runs cited an outdated policy page; evidence—the policy page contains an old date and lacks the current exception; next action—policy owner updates the canonical page and redirects or annotates the obsolete page; uncertainty—no conclusion that the update will change future citations. This form makes the work useful even when engine outputs fluctuate.

Prioritize high-confidence repairs that improve the public answer itself: a wrong date, missing eligibility condition, broken link, blocked documentation page, or unsupported comparison claim. Treat correlations as investigations. A competitor’s repeated presence may justify reviewing equivalent evidence coverage, but it does not prove that copying its heading or issuing more prompts will reproduce the result. Close the weekly loop by noting completed edits, deferred work, source changes, and any reason the next sample will differ. That modest decision log is more valuable to a small team than a precise-looking score with no trail back to the underlying observations.

Related reading: Seven Metrics That Make AI Visibility Measurable.

Methodology

This is a manual audit workflow, not a benchmark. A single prompt run is an observation and is not claimed to be statistically representative.

Sources

  1. Google Search CentralAI Features and Your Website(opens in a new tab)
  2. Google Search CentralRobots Meta Tag Specifications(opens in a new tab)
  3. OpenAI Help CenterAdvertiser Guidance for Allowing OpenAI Web Crawlers(opens in a new tab)
  4. Perplexity DocumentationPerplexity Crawlers(opens in a new tab)
  5. arXivRetrieval-Augmented Generation for Knowledge-Intensive NLP Tasks(opens in a new tab)