On this page

The panel exists to answer one question honestly: when a real buyer works through this decision with an AI assistant, does the path lead toward the brand, and is the brand described accurately when it appears? Everything in the method serves that question, and everything that biases it, especially editing prompts after seeing results, destroys it.

That Google guidance is about content, but it strengthens the audit method too: use related wording to understand one decision, not to manufacture thirty supposedly different intents.

Derive prompts from one buyer decision

Start with a single sentence that names the person, their context, the alternatives, and the decision:

An operations manager at a US HVAC contractor with 20 field technicians is deciding whether to buy dedicated field service management software, extend the accounting suite the company already pays for, or keep dispatching from spreadsheets.

This sentence does the real work. It fixes geography and audience, names the competing options (including the do-nothing option, which is usually the strongest competitor), and gives every prompt a purpose. A prompt either samples a question this buyer would plausibly work through, or it does not belong in the panel.

Do not begin with “what should we ask ChatGPT?” Platform choice comes after the buyer model. The same decision panel is then observed on whichever surfaces you declare, and the 47-point audit checklist requires those declarations before the first run.

The four strata and what each one tests

Problem discovery prompts describe the costly situation without naming a category or brand. They test whether the buyer’s problem, phrased the way an operator phrases it, leads toward the relevant solution space at all.

Category discovery prompts ask what kind of product, provider, or method helps. These are unbranded and commercially decisive: if the category answer omits the brand’s segment or misdefines the category, nothing downstream recovers it.

Comparison and fit prompts test alternatives, tradeoffs, and non-fit conditions. They are worth more than endless “best X” variants because real buyers compare against the specific alternative they already have.

Validation and risk prompts ask whether a named option is credible, supported, and appropriate. Branded prompts live here, and their results must never be blended with discovery results, because a brand that dominates its own name while vanishing from discovery has a serious problem that a blended number would hide.

Prompt panel planner

Start with one balanced panel, then repeat each prompt three times on each declared platform. This estimates observation work, not audience demand.

Problem discovery

9 prompts

Category discovery

9 prompts

Comparison and fit

8 prompts

Validation and risk

5 prompts

Three repeated runs create 270 saved observations across the panel. Add review time for citations, representation, and evidence capture.

The default 30-prompt panel splits roughly 9, 9, 8, and 4 across the strata; the planner’s per-stratum rounding can shift a prompt between adjacent strata. The exact split is not a law; the balance is the point. A panel that is all comparisons measures shortlist behavior and nothing upstream of it.

A complete worked 30-prompt panel

Worked example

A 30-prompt panel for one declared buyer decision

The scenario is the buyer decision sentence above: field service management software evaluated by a US HVAC contractor. The product category is real; the scenario is an illustration, so branded prompts use [brand] and [competitor] placeholders rather than invented company names. Substitute the audited brand and its named competitors.

Problem discovery (9):

IDPrompt
P01How should an HVAC contractor stop losing track of service calls during peak season?
P02What should a contractor do when dispatchers schedule jobs from a whiteboard and spreadsheets?
P03Why do technicians miss follow-up maintenance visits and how do you prevent that?
P04How can a 20-technician HVAC company cut time spent on paper work orders?
P05What causes double-booked service appointments and how do contractors fix it?
P06How should a home services company handle after-hours emergency call scheduling?
P07What is the fastest way for a contractor to send the invoice the same day as the job?
P08How do HVAC companies keep customers informed about technician arrival times?
P09How should a contractor track maintenance agreement renewals so they stop lapsing?

Category discovery (9):

IDPrompt
P10What kind of software helps HVAC contractors schedule and dispatch technicians?
P11What is field service management software and who actually needs it?
P12When should a contractor switch from spreadsheets to field service management software?
P13What features matter most in field service software for residential HVAC work?
P14Does a 20-technician contractor need dedicated dispatch software or is an accounting suite add-on enough?
P15What does field service management software typically cost for a small contractor?
P16What should an HVAC company look for in scheduling software with technician GPS tracking?
P17Can field service software handle maintenance agreements and recurring visit scheduling?
P18What are the risks of running an HVAC business without dispatch software?

Comparison and fit (8):

IDPrompt
P19Field service management software vs an accounting suite scheduling add-on for an HVAC contractor: which fits better?
P20[brand] vs [competitor] for a residential HVAC company
P21Which field service platforms work best for contractors with 10 to 50 technicians?
P22What are the best alternatives to [competitor] for HVAC dispatching?
P23Is an all-in-one field service platform better than separate scheduling and invoicing tools?
P24Which field service software integrates well with existing accounting systems?
P25When is enterprise field service software too heavy for a small contractor?
P26What should a contractor compare before choosing between two field service platforms?

Validation and risk (4):

IDPrompt
P27Is [brand] a reliable choice for a US HVAC contractor?
P28What do reviews say about [brand] customer support and onboarding?
P29What are the limitations of [brand] for maintenance agreement billing?
P30How hard is it to migrate from spreadsheets to [brand]?

Notice what the panel does not contain: ten punctuation variants of the same question, prompts engineered to make [brand] the only sensible answer, and prompts no operator would ever type. P25 exists specifically to test whether non-fit conditions are represented honestly, because a panel with no unflattering prompts is a sales asset, not a sample.

Prompt-set worksheet with the worked panel

All 30 illustration prompts with stratum and the decision each one tests, plus empty capture columns for your own prompts, platform, account state, run number, mention classification, cited URLs, and representation results.

CSV worksheet

Freeze the panel, then control every run

Freeze the panel, with a version number and a date, before the first saved observation. This is the baseline discipline: the first complete set of runs against the frozen panel is the only baseline you will ever get, and every future wave is comparable only if the panel, platforms, account state, location context, and run counts match it.

For each observation, save the prompt ID and exact text, platform and surface, signed-in or signed-out state, date and time, location and language context, run number, the full answer evidence, every cited URL, the mention classification, and the representation result against the approved fact ledger. Run the frozen panel more than once per platform, because generative answers vary between identical runs. Three runs is a practical directional convention, not a statistical law; state the count and its limits rather than calling any of it representative, as covered in the measurement protocol.

When the business genuinely changes, version the panel. If v2 replaces five prompts after the client launches a new offer, keep the unchanged 25 as a continuity core and never draw one trend line across the break.

Keep denominators and classifications honest

Two bookkeeping rules protect the client from the numbers.

First, mentions and citations are different events. A brand named in prose was recalled or synthesized; a brand cited as a linked source was retrieved and used. Count them in separate columns, always.

Second, report strata separately before any total. If the brand appears in 6 of 30 prompts, that is a 20 percent mention rate on that declared panel, nothing more. The useful decomposition is usually: unbranded discovery rate (say 1 of 20, 5 percent), branded validation rate (5 of 10, 50 percent). The blended 20 percent hides exactly the finding that matters, which is that buyers who have not heard of the brand never encounter it. Reporting layers separately is the same principle that structures the four-layer visibility model.

Six of the 47 checks grade the panel on its own

Confirm the derivation, the strata, the freeze, the run counts and the separated denominators, and see whether the sample section holds before the rest of the protocol is graded against it.

Score the prompt sample

Write the buyer-decision sentence, build the four strata, freeze the panel, and only then open an answer engine. Score the whole protocol against the 47-point audit checklist before any finding reaches a client.