What each tool class actually observes
Every vendor in this market sells a window, not the weather. There are four windows, and they answer different questions.
Prompt trackers (Otterly.AI, Profound, Semrush’s AI Visibility Toolkit) run a list of prompts against AI engines on a schedule and record the answers. They observe exactly what your panel asks, under the vendor’s collection conditions, and nothing else. This is the closest match to an audit workflow built on a frozen prompt set, like the one described in how to build an AI visibility prompt set.
Corpus monitors (Ahrefs Brand Radar) mine a large index of search-derived prompts, 405 million or more by Ahrefs’ own description, and report where a brand appears across it. They observe breadth you could never sample manually, but the prompts are the vendor’s, not your buyers’, so treat the output as market context rather than a client baseline.
Referral analytics (GA4 source filters, Bing Webmaster Tools) observe outcomes: humans who clicked from an AI surface to the site. They see nothing about answers where no click happened, which is most of them.
Log analysis (server or CDN logs checked against published crawler IP ranges) observes access: which AI agents actually fetched which pages. It is the only class that can verify the eligibility layer covered in OAI-SearchBot vs GPTBot.
No class observes why a model chose a source, what any individual user saw, or the effect of personalization and memory. A vendor score that implies otherwise is compressing a sample into a certainty.
Decide the criteria before opening a vendor tab
Comparing feature grids first is how agencies end up owning three overlapping subscriptions. Write down the scope you must serve, then test candidates against these criteria in order:
- Method disclosure. How are prompts sourced, where are answers collected (consumer product, API, headless browser), how often, from what location and account state? If the vendor will not say, you cannot defend the numbers to a client.
- Raw evidence access. Can you read and export the full answer text, cited URLs, and timestamps per run? Scores summarize; evidence survives client scrutiny.
- Capacity fit. Does the prompt allowance cover your actual panel across your actual engines, or does the real scope live in add-ons?
- True scope cost. Price one representative client end to end, including engines, markets, seats, and workspaces.
- Workflow fit. Client isolation, report editing, and export into your reporting template.
- Exit path. Historical data portability, and a method you could continue without the tool.
Four vendors compared, checked August 7, 2026
All figures below are vendor-stated, read from public pricing and documentation pages on the check date. They change often; verify on the purchase date.
| Tool | Entry price | Tracked prompts at entry | Engines | Collection approach |
|---|---|---|---|---|
| Otterly.AI | $29/mo (Lite) | 15 | ChatGPT, Google AI Overviews, Perplexity, Copilot; Claude, AI Mode, Gemini as paid add-ons | Scheduled runs of your prompts |
| Semrush AI Visibility Toolkit | $99/mo per domain, billed annually | 25 custom prompts, daily | ChatGPT, Google AI, Gemini, Perplexity | Scheduled runs plus domain-level reports |
| Profound | $99/mo (Starter, billed yearly) | 50 prompts, 1,500 responses/mo | ChatGPT only at Starter; Growth at $399/mo adds Perplexity and Google AI Overviews; Enterprise up to 9 engines | Scheduled runs at volume |
| Ahrefs Brand Radar | Included with Ahrefs plans for custom prompts; standalone $199/mo per platform or $699/mo all platforms | Corpus of 405M+ search-backed prompts; custom prompt checks from $50/mo | Google AI Overviews and AI Mode, ChatGPT, Perplexity, Gemini, Copilot, Grok, Claude | Prompt index mined from search behavior |
Read the table against your scope, not in the abstract. A solo consultant with one client panel of 20 prompts on two engines has no use for a 405-million-prompt corpus. A five-client agency running 100 prompts daily across four engines will burn analyst hours a $189 plan would erase.
When a spreadsheet beats a subscription
The manual method is not the budget option; it is the method-control option. You choose the prompts, conditions, and classification rules, and you keep every answer as evidence. The tradeoff is capture labor.
Worked example
Manual capture cost vs a prompt tracker
Declared assumptions: a 25-prompt client panel, 3 repeated runs per prompt, 2 engines, monthly cadence, 2 minutes to run and log each observation, analyst loaded cost $70/hour.
Observations per month: 25 x 3 x 2 = 150.
Capture time: 150 x 2 minutes = 5 hours. Add 1 hour for setup drift and rechecks: 6 hours.
Manual capture cost: 6 x $70 = $420 per month, before classification and reporting, which both routes still require.
Tool route: Otterly Standard at $189/mo covers the panel with capacity to spare; Semrush at $99/mo covers exactly 25 prompts on one domain. Either is cheaper than manual capture for this scope, but only if its collection conditions and exports satisfy the audit method. If the tool cannot export raw answers with timestamps, the $420 buys evidence the subscription cannot.
The spreadsheet wins when any of these hold: the engagement is a one-time baseline rather than continuous monitoring; the client demands full answer evidence; you are still validating that clients will pay before committing to annual, per-domain billing; or your panel is under roughly 50 observations per run. The subscription wins on daily cadence, multi-client portfolios, and engine breadth. The full manual protocol is in how to measure AI search visibility.
Score candidates with the worksheet
Run every shortlisted tool through the same 24 criteria before paying. The scorecard below implements the criteria from this article with weights, a 0 to 2 scoring scale, and space for evidence notes, so two people evaluating the same tool produce comparable results.
AI visibility tool evaluation scorecard24 weighted criteria across observation method, evidence access, coverage and cost, workflow fit, and exit path, with a question to ask, passing evidence, and scoring columns for each.
CSV worksheet Score the method before you score the vendorsThe scorecard above rates what a tool exposes. This one rates what you do with it, across all 47 checks, so you can tell a capability you are missing from a subscription that would not have fixed it.
Grade your own protocolOnce the tool decision is made, fold its outputs into the audit workflow: the 47-point audit checklist covers what to inspect, and GEO vs SEO frames where tool-based observation sits inside the wider engagement. The AI Search Visibility Audit System includes the full workbook the scorecard feeds into.