The panel exists to answer one question honestly: when a real buyer works through this decision with an AI assistant, does the path lead toward the brand, and is the brand described accurately when it appears? Everything in the method serves that question, and everything that biases it, especially editing prompts after seeing results, destroys it.
That Google guidance is about content, but it strengthens the audit method too: use related wording to understand one decision, not to manufacture thirty supposedly different intents.
Derive prompts from one buyer decision
Start with a single sentence that names the person, their context, the alternatives, and the decision:
An operations manager at a US HVAC contractor with 20 field technicians is deciding whether to buy dedicated field service management software, extend the accounting suite the company already pays for, or keep dispatching from spreadsheets.
This sentence does the real work. It fixes geography and audience, names the competing options (including the do-nothing option, which is usually the strongest competitor), and gives every prompt a purpose. A prompt either samples a question this buyer would plausibly work through, or it does not belong in the panel.
Do not begin with “what should we ask ChatGPT?” Platform choice comes after the buyer model. The same decision panel is then observed on whichever surfaces you declare, and the 47-point audit checklist requires those declarations before the first run.
The four strata and what each one tests
Problem discovery prompts describe the costly situation without naming a category or brand. They test whether the buyer’s problem, phrased the way an operator phrases it, leads toward the relevant solution space at all.
Category discovery prompts ask what kind of product, provider, or method helps. These are unbranded and commercially decisive: if the category answer omits the brand’s segment or misdefines the category, nothing downstream recovers it.
Comparison and fit prompts test alternatives, tradeoffs, and non-fit conditions. They are worth more than endless “best X” variants because real buyers compare against the specific alternative they already have.
Validation and risk prompts ask whether a named option is credible, supported, and appropriate. Branded prompts live here, and their results must never be blended with discovery results, because a brand that dominates its own name while vanishing from discovery has a serious problem that a blended number would hide.
Prompt panel planner
Start with one balanced panel, then repeat each prompt three times on each declared platform. This estimates observation work, not audience demand.
Problem discovery
9 prompts
Category discovery
9 prompts
Comparison and fit
8 prompts
Validation and risk
5 prompts
The default 30-prompt panel splits roughly 9, 9, 8, and 4 across the strata; the planner’s per-stratum rounding can shift a prompt between adjacent strata. The exact split is not a law; the balance is the point. A panel that is all comparisons measures shortlist behavior and nothing upstream of it.
A complete worked 30-prompt panel
Worked example
A 30-prompt panel for one declared buyer decision
The scenario is the buyer decision sentence above: field service management software evaluated by a US HVAC contractor. The product category is real; the scenario is an illustration, so branded prompts use [brand] and [competitor] placeholders rather than invented company names. Substitute the audited brand and its named competitors.
Problem discovery (9):
| ID | Prompt |
|---|---|
| P01 | How should an HVAC contractor stop losing track of service calls during peak season? |
| P02 | What should a contractor do when dispatchers schedule jobs from a whiteboard and spreadsheets? |
| P03 | Why do technicians miss follow-up maintenance visits and how do you prevent that? |
| P04 | How can a 20-technician HVAC company cut time spent on paper work orders? |
| P05 | What causes double-booked service appointments and how do contractors fix it? |
| P06 | How should a home services company handle after-hours emergency call scheduling? |
| P07 | What is the fastest way for a contractor to send the invoice the same day as the job? |
| P08 | How do HVAC companies keep customers informed about technician arrival times? |
| P09 | How should a contractor track maintenance agreement renewals so they stop lapsing? |
Category discovery (9):
| ID | Prompt |
|---|---|
| P10 | What kind of software helps HVAC contractors schedule and dispatch technicians? |
| P11 | What is field service management software and who actually needs it? |
| P12 | When should a contractor switch from spreadsheets to field service management software? |
| P13 | What features matter most in field service software for residential HVAC work? |
| P14 | Does a 20-technician contractor need dedicated dispatch software or is an accounting suite add-on enough? |
| P15 | What does field service management software typically cost for a small contractor? |
| P16 | What should an HVAC company look for in scheduling software with technician GPS tracking? |
| P17 | Can field service software handle maintenance agreements and recurring visit scheduling? |
| P18 | What are the risks of running an HVAC business without dispatch software? |
Comparison and fit (8):
| ID | Prompt |
|---|---|
| P19 | Field service management software vs an accounting suite scheduling add-on for an HVAC contractor: which fits better? |
| P20 | [brand] vs [competitor] for a residential HVAC company |
| P21 | Which field service platforms work best for contractors with 10 to 50 technicians? |
| P22 | What are the best alternatives to [competitor] for HVAC dispatching? |
| P23 | Is an all-in-one field service platform better than separate scheduling and invoicing tools? |
| P24 | Which field service software integrates well with existing accounting systems? |
| P25 | When is enterprise field service software too heavy for a small contractor? |
| P26 | What should a contractor compare before choosing between two field service platforms? |
Validation and risk (4):
| ID | Prompt |
|---|---|
| P27 | Is [brand] a reliable choice for a US HVAC contractor? |
| P28 | What do reviews say about [brand] customer support and onboarding? |
| P29 | What are the limitations of [brand] for maintenance agreement billing? |
| P30 | How hard is it to migrate from spreadsheets to [brand]? |
Notice what the panel does not contain: ten punctuation variants of the same question, prompts engineered to make [brand] the only sensible answer, and prompts no operator would ever type. P25 exists specifically to test whether non-fit conditions are represented honestly, because a panel with no unflattering prompts is a sales asset, not a sample.
All 30 illustration prompts with stratum and the decision each one tests, plus empty capture columns for your own prompts, platform, account state, run number, mention classification, cited URLs, and representation results.
CSV worksheetFreeze the panel, then control every run
Freeze the panel, with a version number and a date, before the first saved observation. This is the baseline discipline: the first complete set of runs against the frozen panel is the only baseline you will ever get, and every future wave is comparable only if the panel, platforms, account state, location context, and run counts match it.
For each observation, save the prompt ID and exact text, platform and surface, signed-in or signed-out state, date and time, location and language context, run number, the full answer evidence, every cited URL, the mention classification, and the representation result against the approved fact ledger. Run the frozen panel more than once per platform, because generative answers vary between identical runs. Three runs is a practical directional convention, not a statistical law; state the count and its limits rather than calling any of it representative, as covered in the measurement protocol.
When the business genuinely changes, version the panel. If v2 replaces five prompts after the client launches a new offer, keep the unchanged 25 as a continuity core and never draw one trend line across the break.
Keep denominators and classifications honest
Two bookkeeping rules protect the client from the numbers.
First, mentions and citations are different events. A brand named in prose was recalled or synthesized; a brand cited as a linked source was retrieved and used. Count them in separate columns, always.
Second, report strata separately before any total. If the brand appears in 6 of 30 prompts, that is a 20 percent mention rate on that declared panel, nothing more. The useful decomposition is usually: unbranded discovery rate (say 1 of 20, 5 percent), branded validation rate (5 of 10, 50 percent). The blended 20 percent hides exactly the finding that matters, which is that buyers who have not heard of the brand never encounter it. Reporting layers separately is the same principle that structures the four-layer visibility model.
Six of the 47 checks grade the panel on its ownConfirm the derivation, the strata, the freeze, the run counts and the separated denominators, and see whether the sample section holds before the rest of the protocol is graded against it.
Score the prompt sampleWrite the buyer-decision sentence, build the four strata, freeze the panel, and only then open an answer engine. Score the whole protocol against the 47-point audit checklist before any finding reaches a client.