On this page

Use the checklist at three moments: before fieldwork to catch missing scope, during the audit to keep the evidence record complete, and before handoff to reject findings that nothing supports. An audit that skips the third pass ships opinions with screenshots attached.

How the 47-point scoring model works

Every check carries a weight of two or three. Weight-three checks, 29 of the 47, mark conditions that can invalidate a conclusion on their own: missing consent, an unfrozen prompt set, a finding with no saved evidence. Weight-two checks mark gaps that weaken the audit without collapsing it. The score is earned weighted points divided by 123 possible points, times 100.

The client-ready verdict has two gates, not one. The weighted score must reach 85, and all 29 weight-three checks must be confirmed. A missing critical check caps the verdict at Usable with material gaps, even when the arithmetic score is 85 or higher. Weight-two gaps may remain only when the score still reaches 85 and a human reviewer judges them acceptable for the declared scope.

Score and critical-check conditionInterpretationAllowed use
85 to 100, with every weight-three check confirmedClient-ready protocolFindings can enter a client report after human review
85 to 100, with any weight-three check still missingUsable with material gapsClose every critical gap before a final recommendation
65 to 84Usable with material gapsClose the named gaps before a final recommendation
40 to 64Directional audit onlyInternal discovery, never outcome claims
Below 40Not yet defensibleRe-scope and rebuild the evidence record

A high score does not mean the brand is visible. It means the audit can explain its own evidence. Those are different achievements, and only the second one is under your control.

Score your own protocol against all 47 checks

The interactive version of this model. Confirm what you already record, see which sections are weak, and get every unconfirmed check ranked by weight with the evidence to keep and the failure that costs the point.

Open the checker
The 47-point audit checklist as a working file

All 47 checks with section, weight, the evidence artifact to save for each check, and the failure that most often costs the point. Use it as the quality tab of the audit workbook this model grades.

CSV template

The first eleven checks exist because most bad audits fail before anyone opens an answer engine. Scope failures are quiet: nobody wrote down the one business decision being tested, so the prompt panel drifts toward whatever wording produces interesting screenshots. Consent failures are louder. Auditing a prospect’s brand without recorded permission produces evidence you cannot present, and in the worst case a sales deck built on observations the prospect never authorized.

The fact ledger is the check practitioners skip most. Representation review, the part of the audit clients care about most, is impossible without it. If the client has not approved, in writing, the canonical business name, the offers and their fit conditions, the locations, and the sources behind every credential claim, then the auditor grades AI answers against their own guesses. When an engine repeats an old service area, an audit without a dated ledger cannot even say whether that is an error.

Weight-three checks here are consent, approved offers, sourced claims, and written exclusions. The exclusions check surprises people: it requires that what the audit does not promise (rankings, citations, traffic, revenue) exists in writing before fieldwork. That single artifact prevents the most expensive handoff argument, and it is why the audit proposal carries a dedicated exclusions section.

Technical eligibility before any prompt runs

The seven technical checks come third for a reason of arithmetic, not ideology. Google’s public position is that AI Overviews and AI Mode have no special technical requirements beyond normal indexing and snippet eligibility (Google Search Central), and ChatGPT search inclusion runs through an ordinary crawler allowance, as covered in OAI-SearchBot vs GPTBot. So if a priority page returns the wrong public response, blocks the relevant crawler, carries a stray noindex, or hides its material facts in client-side rendering, every downstream observation about that page is noise. You would be measuring the visibility of a page the systems cannot fully use.

What fails in practice, in rough order of frequency: a blanket bot block added years ago during a scraping scare that now also blocks search crawlers; material facts that exist only in JavaScript-rendered widgets or images; structured data that disagrees with visible copy on price, address, or offer names; and technical checks performed while logged in to a CDN-whitelisted network, so the auditor never sees the challenge page the public gets. The seventh check, dated evidence for every technical claim, exists because “robots is fine” without a date is unfalsifiable one site release later.

The prompt sample and the observation log

Thirteen checks cover sampling and observation because this is where audits most often stop being audits. The prompt checks enforce one discipline above all: the panel is derived from buyer decisions, then frozen before the first saved run. Tuning wording after seeing which phrasing favors the client converts the audit into a demonstration. The full derivation method, including the four strata and the branded-versus-unbranded separation, is in the prompt set guide.

The observation checks are storage discipline. Exact prompt text, full answer evidence, run date, platform, surface, and account state, for every run, including runs where the brand was absent. Absence is data. The two classification checks are weighted three because their failure produces the most misleading client numbers: counting a prose mention and a cited link as the same event overstates source use, and skipping the representation check against the fact ledger means the audit reports presence without ever asking whether what was said is true. Repeated runs matter because generative answers vary; a single run reported as “the state of ChatGPT” is the sampling equivalent of polling one person.

Sources, findings, and the client handoff

The last sixteen checks turn a pile of observations into something a client can act on without being misled. The source map checks force two separations: owned versus third-party cited sources (so the roadmap targets pages someone can actually edit), and competitor numbers computed on the identical frozen panel. A competitor comparison built on different prompts, dates, or run counts is fiction with a table format.

The findings checks encode the difference between observation and explanation. “The brand was absent from 14 of 20 unbranded discovery prompts” is an observation. “The engine penalizes the site” is an invented mechanism nobody observed, and check source-5 requires it to stay labeled unknown. The handoff checks then protect the client from your own report: method limits visible in the main document rather than an appendix, the four evidence layers of the four-layer visibility model kept separate rather than blended into one proprietary percentage, and a roadmap whose items would still be worth doing if no sampled answer ever changed. That last check is weighted three because it is the honesty test for the entire engagement.

A passing executive finding reads like this: across the declared 30-prompt panel and three repeated runs, the firm appeared in category answers, but its US service area was omitted in four of seven mentions, and the approved ledger disagrees with two public profiles on geography. Correct the profiles, publish one canonical fact source, and repeat the identical panel after the changes are discoverable. Sample, observation, factual basis, action, review method. No mechanism claims, no promises.

What to refuse to score

If a client asks for these numbers anyway, the honest substitutes exist: mention and citation rates over a declared panel with the denominator shown, tracked over identical reruns, as described in the measurement protocol.

Rerunning the audit without breaking the baseline

The checklist is also the re-audit instrument. A quarterly or half-yearly wave only produces a trend if the panel, platforms, run counts, account state, and classification rules match the baseline exactly, which is what check handoff-6 locks in at delivery time. When the business genuinely changes (new offer, new geography), version the panel, keep a stable core of unchanged prompts for continuity, and label the break in every chart. Small movements between two waves of a 30-prompt panel are usually variance, not trend; treat single-digit swings as noise unless they persist across waves.

Score the protocol before the first client sees a finding, price the scope you actually declared using the pricing model, and hand the completed evidence record to the client report template. In that order. The report is the last step because everything defensible about it is created earlier.