The complete ten-section structure with placeholder fields, per-section guidance notes, the finding block with separated observation, interpretation, and recommendation fields, and the metrics tables with denominator columns built in.
Markdown templateWhat the report has to survive
A client reading their first AI visibility report brings three reasonable questions: what is actually happening, how do you know, and what should we do about it. The report also has to survive a fourth reader you will never meet: the skeptical colleague, procurement reviewer, or replacement agency who inherits it later and checks whether the numbers trace to anything.
Rank-tracking reports survived on familiarity. AI visibility reports cannot, because the evidence is unfamiliar: sampled answers instead of positions, classification instead of counting, limits that are structural rather than apologetic. So the template’s job is to make unfamiliar evidence legible without pretending it behaves like rank tracking. That drives every structural choice below: short and decisive at the front, fully traceable at the back, method visible in the middle because the method is part of what the client is paying for.
Everything in the report is built on the numbers produced by the measurement protocol. If the observation log does not exist, there is nothing to report, and no template fixes that.
Sections 1 to 3: the front the client reads
Section 1, decision and scope, names the business decision the report supports, then the frozen parameters: audience, geography, platforms, the prompt panel size and strata, run counts, priority pages, and explicit exclusions. The reasoning: a report without a named decision gets read as a score, and scores get argued with. A report scoped to a decision gets used.
Section 2, executive finding, is one page answering the three client questions in order: what we observed (one to three findings, each with numerator and denominator), how we know (two sentences pointing at the method), and the recommended next action with an owner and an acceptance check. Write it last, compress it hardest.
Section 3, method and limits, states that the report is a controlled sample of observed answers, not a census of AI answers or users; how prompts were selected; what account state and personalization applied; and that nothing in the report establishes a ranking or predicts future mentions, citations, traffic, or revenue. Do not shrink this into a footnote. When a competitor’s report waves a black-box score around, your visible method is the differentiation.
Sections 4 to 7: the evidence core
Section 4, technical eligibility, reports first-party checks on priority pages: response codes, robots and index controls, canonicals, and whether material facts appear in rendered text, each with a checked date. It stays separate from answer observations because the two evidence types fail independently: a page can pass every check and never be cited, and a cited page can carry a representation error. This section is the first layer of the four-layer model, and the report keeps all four layers in separate sections for the same reason the model does.
Section 5, mentions and citations, is the baseline table: valid observations, brand mentions, owned citations, and their rates, split by platform and by prompt stratum. Branded and unbranded strata never share a denominator; a pooled rate flatters recognition and hides the discovery problem.
Section 6, representation accuracy, compares each material statement in sampled answers against the approved fact ledger and labels it accurate, incomplete, outdated, conflicting, or unsupported. This is usually where the commercially urgent findings live, because a confident wrong answer in front of a buyer outranks any visibility statistic.
Section 7, sources and competitors, maps which owned and third-party sources appeared, which sources answered the buyer’s questions without the brand, and which competitors were named per stratum. This is the section that converts observation into a work plan, because it shows where the evidence gaps physically are.
First-party platform data joins the report here as its own labeled subsection, never merged into panel rates: Bing’s AI Performance preview reports aggregated citation activity with its own coverage and caveats (the reconciliation workflow covers the details), and ChatGPT referral clicks are attributable through the utm_source=chatgpt.com parameter documented in OpenAI’s publisher FAQ.
Sections 8 to 10: findings, roadmap, appendix
Section 8, findings and priorities, is one block per finding, each preserving the same chain: observation, interpretation, buyer consequence, recommendation, owner, confidence, acceptance check. The chain is the integrity mechanism; the worked example below shows why the first three fields must never blur.
Section 9, roadmap, sequences actions by dependency, not effort: eligibility blockers before content work, factual conflicts before anything that would amplify them. Each row names the finding it derives from. A roadmap item with no parent finding is an upsell, and clients can tell.
Section 10, appendix, carries prompt IDs, the observation log reference, the fact ledger version, definitions, and a change log. Raw answer dumps stay in the evidence workbook; the appendix points to them.
Writing one finding: the three-field discipline
Worked example
One finding, written as observation, interpretation, and recommendation
This illustration uses abstract placeholders throughout. It describes no real or invented client, and its numbers exist only to show the structure.
Observation. In the [period] panel, [Brand]‘s service pages were cited in 0 of 36 valid unbranded observations on [platform]. The eligibility check found the service directory returns its content only after client-side rendering, and the rendered-text check failed for 4 of 6 priority pages (evidence rows [IDs], checked [date]).
Interpretation. We assess, labeled as inference, that the rendering gap is the earliest plausible blocker: the platform does not expose why any page was or was not cited, but pages whose material facts are absent from rendered text cannot be expected to appear as sources. Competing explanations, thin third-party evidence among them, remain open until the blocker is cleared.
Recommendation. Ship server-rendered content for the 4 failing pages (owner: [role]), verify with the same rendered-text check, then rerun the identical panel after the pages are re-crawlable. Acceptance check: rendered-text pass on 6 of 6 priority pages; the citation rate itself is not a promise and is not the acceptance criterion.
The discipline: the observation contains only what was recorded. The interpretation is labeled as inference and admits alternatives. The recommendation is the smallest action with an acceptance check the agency actually controls.
Cadence, versioning, and what to cut
Report on a cadence the noise floor can support. A monthly full panel is the workable default for most engagements; weekly reporting of a 30-prompt panel mostly reports variance, and a quarterly gap leaves representation errors live too long. Between full waves, report only first-party data and shipped work. Whatever the cadence, version the method: any change to prompts, platforms, run counts, or classification rules gets a new method version in the report header and starts a new comparable series. A trend across method versions is a fabrication with extra steps.
Cut without mercy: decorative screenshots that add no evidence, proprietary scores with no formula, raw prompt dumps in the client layer, generic GEO education, and any recommendation without a parent finding. Before the report ships, run the whole package through the 47-point audit checklist; findings that fail its evidence checks go back to the workbook, not to the client.
Six handoff checks decide whether this report is deliverableMethod limits in the main document, the four layers kept separate, a roadmap that survives without a citation, and a rerun spec. Confirm what you already do and the checker ranks the rest by weight.
Score the handoff