Short answer, as of August 2026: build 20 to 40 prompts drawn from language your customers actually use, spread across four categories — problem, comparison, recommendation and brand — and sample each one at least five times per engine per cycle. Anything less than that repetition measures noise; anything much more than forty prompts costs money without changing the decision you're about to make.
The prompt set is the whole instrument. Software that samples it is a commodity; deciding what to sample is the judgment call, and it's the one most people skip.
Why keyword research is the wrong starting point
Keyword lists are compressed. Someone wanting a plumber types "emergency plumber near me" because search engines rewarded terse phrasing for twenty years.
The same person asks an assistant: "My kitchen sink is backing up and there's water on the floor — who should I call in Springfield and roughly what does this cost on a Sunday?"
That's one prompt containing a symptom, a location, an urgency signal and a price question. Exporting your top keywords and appending question marks produces a set that looks reasonable and tests language nobody uses. Prompts are longer, messier, more contextual and more likely to carry constraints.
Start instead from records of how people actually describe their problem before they know your vocabulary.
Where the good prompts come from
In rough order of value:
1. Your own inbound. Sales call recordings, chat transcripts, support tickets, contact-form free text. This is real language from real buyers and it's the highest-quality source you have. Read fifty and you'll have thirty prompts.
2. Search Console queries with question intent. Filter for queries containing who, what, how, why, best, near, vs, cost. These are conventional searches, but they map onto conversational phrasing more directly than head terms do.
3. The assistants themselves. Ask ChatGPT what someone in your situation typically wants to know before hiring in your category. Models are good at generating the space of plausible questions, since it's a paraphrase task.
4. Community sources. Reddit threads, industry forums, review text. Reviews are especially useful because they state what mattered after the purchase, which is usually what the next buyer should have asked.
5. Your competitors' FAQ pages. They've already done this exercise. Their FAQ headings are a free list of what buyers ask.
What to skip: pure head terms ("plumber Springfield"), anything with no purchase relationship, and prompts so specific only one person would ever type them.
The four categories, and why you need all of them
A set weighted entirely to one category produces a chart that moves for reasons you can't act on.
| Category | Example shape | What it measures | Share of set |
|---|---|---|---|
| Problem | "My X is doing Y — what's wrong and who fixes it?" | Whether you're surfaced at the moment of need | ~40% |
| Comparison | "Is A or B better for someone who needs C?" | Whether you appear in evaluation | ~25% |
| Recommendation | "Who's the best provider of X in Y, with good reviews?" | Direct competitive share of voice | ~25% |
| Brand | "Is [your company] any good? What do they charge?" | What engines say when someone checks you out | ~10% |
The brand category is the one most often left out and the one that most often produces an emergency. It tells you what an engine says about you to a prospect doing diligence — and that answer can be wrong, stale, or drawn from a competitor's comparison page. You want to know before a customer tells you.
Recommendation prompts are the closest thing to a competitive scoreboard, since they force the engine to name specific businesses. Problem prompts are the leading indicator: they're where a well-executed content programme shows up first, because a page answering a specific symptom is retrievable in a way a brand page isn't.
How many, and how many times
Twenty to forty prompts for a single-location business or a focused product. Enough to cover the categories, small enough to sample properly on a modest budget. Multi-location or multi-product operations need more — but scale by distinct buying situations, not by page count. Eight locations selling the same service to the same buyer is one prompt set with a location variable, not eight sets.
At least five samples per prompt per engine per cycle. This is not optional and it's where most homemade tracking falls apart. Generated answers are non-deterministic: the same prompt run twice can produce different businesses. Sampled once, you've measured a coin flip and drawn a trend line through randomness. Five samples gives a usable rate; ten is better when the decision is expensive.
Six engines, because coverage differs sharply between them. Being named consistently in Perplexity and never in Gemini is a real and common pattern, and it's actionable — the engines weight sources differently, as covered in how ChatGPT, Perplexity and Gemini pick their sources. One engine's number is not "AI visibility."
Multiply it out: 30 prompts × 6 engines × 5 samples = 900 checks a cycle. That arithmetic is why this belongs in software and why plans are priced in prompt-checks per month.
Writing prompts that behave
A few mechanical rules that make a set produce clean data:
Write like a person, not a query. "Best CRM for a 12-person field service team that already uses QuickBooks" beats "best field service CRM."
One question per prompt. Compound prompts produce compound answers and you can't tell which half moved.
Include the constraint that matters. Location, budget, timeline, industry — whatever genuinely narrows the answer. Prompts without constraints return generic national brands and teach you nothing about your market.
Freeze the wording. Once a prompt is in the set, don't edit it. Rewriting mid-quarter breaks the time series, and you will not remember which change caused which movement.
Avoid prompts that name you outside the brand category. "Why is [your company] the best X?" makes the model agreeable rather than informative.
Add new prompts as additions, not replacements. Keep the original set intact so history stays comparable, and track the new ones as their own cohort.
Reading the output without fooling yourself
Share of voice on a prompt — how often you're named across samples — is the headline number, and it needs care.
A move from 0% to 20% on one prompt is meaningful; 40% to 45% on one prompt is probably noise at five samples. Read the set in aggregate and read categories separately. Watch for prompts where you were never named and now sometimes are: that's the leading edge of the work landing, and it usually shows up in problem prompts first because content moves fastest there.
Also watch who else is named. The competitive set an engine returns is frequently not the one you'd have listed — that's often the most valuable single output of a first baseline. When a name you don't recognise keeps appearing, the diagnostic in what to do when a competitor is cited and you aren't is the next step.
And keep the measurement honest about what it can't see: share of voice moves before traffic does, and much of the resulting demand never carries a referrer at all. Why that's structural, and what to pair it with, is in why AI citations barely show up in analytics.
Turning the set into a work queue
A prompt set is not just a scoreboard. Every prompt where you're absent is a content brief with the audience research already done — you know the exact question, in the customer's own words, that you're failing to answer. Prioritise by commercial value rather than by gap size: being missing from a high-intent recommendation prompt matters more than being missing from a broad informational one.
That's the loop a first quarter should run on, and it's essentially the structure of the first 30 days playbook.
Where the software fits
DigiRank's AI Visibility Tracker takes the set you define, samples it on a schedule across ChatGPT, Claude, Gemini, Perplexity, Grok and Google AI Mode, and charts share of voice per prompt over time — with an opportunities view that lists every prompt you're missing from and a one-click path from a gap to a drafted article. Prompt-checks are the unit plans are sized in: Starter at $99/mo includes 200 checks a month for one brand, Agency at $249/mo includes 2,500 across up to 15 client tenants with a 14-day trial, and Pro at $499/mo raises it to 10,000. Sampling runs alongside Search Console, GA4 and Google Business Profile so a share-of-voice rise and the branded-search echo that follows it sit on one chart. What a sampling run does mechanically is in how an AI search visibility tracker works; how to evaluate vendors is in the tracker buyer's guide.
Spend the afternoon on the prompts. Everything downstream inherits their quality.
Frequently asked questions
How many prompts should I track for AI visibility? Twenty to forty for a single-location business or focused product, spread across problem, comparison, recommendation and brand categories. Scale up by distinct buying situations rather than by number of pages or locations — eight branches selling the same service to the same buyer is one prompt set with a location variable.
How many times should each prompt be sampled? At least five times per engine per cycle, and ten when the decision is expensive. Generated answers are non-deterministic, so a prompt checked once measures randomness rather than visibility. Insufficient repetition is the most common flaw in homemade tracking.
Can I just convert my keyword list into prompts? No — that's the most common mistake. Keywords are compressed by two decades of search-engine habit; prompts are long, contextual and carry constraints like location, budget and urgency. Source prompts from sales calls, support tickets and review language instead.
Which engines should I track? All the ones your buyers plausibly use — typically ChatGPT, Claude, Gemini, Perplexity, Grok and Google AI Mode. Coverage varies sharply between them because they weight sources differently, so a single engine's number is not a measure of AI visibility.
Should I change my prompts when I think of better ones? Add rather than replace. Editing a prompt's wording breaks the time series and you won't be able to tell later whether movement came from your work or from the rewording. Keep the original set frozen and track additions as a separate cohort.
What do I do with prompts I'm never named in? Treat each as a content brief with the audience research already done — you have the exact question in the customer's own words. Prioritise by commercial value rather than by gap size; missing from a high-intent recommendation prompt matters more than missing from a broad informational one.
Should I include prompts that mention my brand name? Yes, but keep them to roughly a tenth of the set and phrase them neutrally, such as "Is [company] any good and what do they charge?" This tells you what an engine says to a prospect doing diligence, which is frequently stale or sourced from a competitor. Avoid leading phrasings that invite agreement.