On this page

What is observable, and what vendors imply

Three kinds of evidence about AI search visibility actually exist.

First, sampled answers. You can run a declared set of prompts on a declared platform on a declared date and record what came back: whether the brand was mentioned, whether an owned page was cited, whether the statements about the brand were accurate. That is a sample, and it only supports claims about itself.

Second, first-party platform reporting. Google states that AI Overviews and AI Mode traffic is counted inside the Performance report’s Web search type, with no separate AI segment, in its AI features documentation. Bing’s AI Performance preview reports aggregated citation activity for supported Microsoft experiences.

Third, referral traffic. OpenAI’s publisher FAQ says ChatGPT appends utm_source=chatgpt.com to outbound links, which makes clicked visits attributable in analytics.

Everything else is inference. Nobody outside the platforms observes the full population of prompts, the personalized answers other users see, or the influence of an answer that never produced a click. Vendor dashboards that print one visibility score are sampling too; they are just not showing you the sample. The four-layer model is the map for keeping these evidence types apart: eligibility and outcomes come from first-party checks and analytics, evidence and representation come from sampled answers.

The protocol: five working days, one spreadsheet

Day 1: freeze the scope and the facts. Write down the business decision, geography, language, platforms, observation dates, priority pages, and exclusions. Then build the approved fact ledger: the brand’s name, offers, locations, credentials, and pricing language as the client wants them stated, each with a source and an owner. Decide now what counts as a material representation error. Scope written after the data arrives is not scope, it is rationalization.

Day 2: freeze the prompt panel. Build 30 prompts from real buyer decisions, split across the four strata of the prompt-set method: roughly 9 problem discovery, 9 category discovery, 8 comparison and fit, and 4 branded validation and risk. That page covers derivation; the rule that matters here is that the panel is frozen before the first run and never edited after you see which wording flatters the client.

Days 3 and 4: run and log. Run every prompt three times per platform, logged out unless the scope says otherwise. One spreadsheet row per prompt per run per platform. Save the full answer text and every cited URL. For a 30-prompt panel on two platforms, that is 180 rows. It is tedious. It is also the entire difference between “I saw” and “we observed.”

Day 5: classify, then calculate. Classification comes first because a rate computed over unclassified rows cannot be audited. The next section defines the fields.

AI visibility observation log

The exact log used in this protocol: one row per observation with prompt stratum, run number, account state, mention and citation fields, the six-value representation label, evidence link, and validity flag, plus a field guide explaining each column.

CSV template

Classify before you calculate

Each row gets independent fields, because the events are independent:

  • brand mentioned: yes or no
  • owned source cited: yes or no
  • third-party source cited: yes or no
  • representation: accurate, incomplete, outdated, conflicting, unsupported, or not applicable
  • competitors named: exact names
  • valid observation: yes or no

A mention is not a citation. A citation is not an endorsement. An accurate mention without a citation and a cited page under an inaccurate summary are different findings that demand different work, and a merged field would erase the difference. The citation side of the log has its own expanded schema, with per-URL records and ownership classification, in the AI citation tracking guide.

The validity flag is the quiet workhorse. A row missing its saved answer, exact prompt, date, or platform is marked invalid and never enters a denominator. If the method allows, rerun it; if not, the denominator shrinks and the report says so.

Rates that publish their denominators

For a 30-prompt panel run three times, the denominator is 90 observations per platform. Compute per platform and per stratum, never pooled:

mention rate = observations with a brand mention ÷ valid observations

owned citation rate = observations citing an owned URL ÷ valid observations

representation error rate = material mentions with an error ÷ material mentions reviewed

competitor appearance rate = observations naming the competitor ÷ valid observations

Branded and unbranded strata must never share a denominator. On a 30-prompt panel run three times, the 4 branded validation prompts yield 12 observations and the 18 unbranded discovery prompts yield 54. A brand that appears in 10 of the 12 branded observations and 2 of the 54 unbranded ones has not “appeared in 12 of 66.” It has a recognition result and a discovery problem, and the pooled number hides both.

Establish a noise floor before claiming change

Answers vary run to run even when nothing about the site changed. Before attributing movement to your work, measure how much the method moves on its own: run the frozen panel in two waves a few days apart, before any changes ship, and use the spread between waves as your practical noise floor.

Worked example

Reading a follow-up wave against a two-wave baseline

All numbers here are declared assumptions to show the arithmetic, not observations from any real brand.

Assume a 30-prompt panel, three runs, one platform: 90 valid observations per wave, of which the 18 unbranded discovery prompts contribute 54. The example tracks the unbranded discovery rate, because that is the number client work usually needs to move.

WaveTimingUnbranded mentionsRate
Baseline AWeek 111 of 5420 percent
Baseline BWeek 2, no changes shipped13 of 5424 percent
Follow-upWeek 8, after fixes15 of 5428 percent

The two baselines differ by 4 points with zero intervention, so 4 points is the observed noise floor of this method. The follow-up sits 4 points above baseline B: inside the floor, so the honest reading is “no detectable change yet.” A follow-up at 23 of 54 would clear the floor by a wide margin and justify investigation, but it would still be a sampled observation, not proof the fixes caused it.

This is a working heuristic, not a confidence interval. Its job is to stop a 3-point swing from becoming a victory slide.

First-party data belongs beside the log, not in it

Use these sources as separate columns of evidence with their own denominators. Bing counts citations across all traffic to supported Microsoft experiences; your panel counts observations of 30 frozen prompts. Both are real, and adding them together produces a number that means nothing. The practical reconciliation workflow is in the Bing AI Performance guide.

Where the numbers go next

Seven of the 47 checks cover the observation log alone

Character-exact prompt text, account state, full answer evidence, absence classified alongside every mention, and citations counted apart from prose. Tick what your spreadsheet already holds and the rest comes back ranked by weight.

Score the observation record

Run the panel through the 47-point audit quality gate before any number reaches a client, then write the findings into the client report template, which keeps observation, interpretation, and recommendation in separate fields. If you are deciding whether a tracking tool should replace the spreadsheet, apply the criteria in the method-first tools comparison: a tool that will not show its prompts, runs, and denominators is asking you to report numbers you cannot defend.