Short answer, as of August 2026: there is no published industry benchmark for AI citation rate, and any vendor quoting one is inventing it. The only benchmark that means anything is your own market: the set of competitors actually named in answers to the questions your buyers ask. Measure your share of that set, date it, and re-measure on a schedule. That comparison is defensible; an absolute score out of 100 is not.
Here is how to build the baseline, what to compare against, and how to read movement without fooling yourself.
Why an absolute score cannot be a benchmark
Three reasons an out-of-context number fails.
Categories differ enormously. A question with three credible national suppliers behaves nothing like one with four hundred local ones. A 40% citation rate might be dominant in the first and unremarkable in the second.
Prompt sets are not comparable. Citation rate is entirely determined by which prompts you chose. Ask ten questions where your brand name appears and you will score wonderfully while learning nothing. Ask ten questions a stranger would actually ask and the number drops and becomes useful. Two companies quoting "62% visibility" may not be measuring anything alike.
Answers are non-deterministic. The same prompt asked twice can produce different sources. One run is a sample, not a measurement, so a single-run score carries an error bar nobody prints.
What survives all three is relative and repeated: of the sources named in answers to these specific questions, what proportion of the naming goes to us versus to the named competitors, measured the same way each time?
The four numbers worth baselining
Before any optimization, record these. It takes a day.
1. Citation rate. Of your prompt-checks, in what percentage were you named? A prompt-check is one prompt asked of one engine once, so a 15-prompt set across six engines is 90 checks. If you appear in 12, that is 13%.
2. Share of voice. Of all brand mentions across those same answers, what proportion are yours? If the answers name 140 brands in total and 12 are you, your share of voice is roughly 9%. This is the number that survives a change in market conditions, because it moves relative to competitors rather than absolutely.
3. Per-engine breakdown. The same four numbers, split across ChatGPT, Claude, Gemini, Perplexity, Copilot and Google AI Mode. Averaging is the most common mistake in this discipline: a jump on one engine and a fall on another cancel to "no change" in an average, when what actually happened is highly informative.
4. The competitor set. Not the competitors you think you have — the ones the engines name. This list is frequently the most valuable output of the first run, and it is routinely surprising. Directories, marketplaces and publishers show up alongside actual rivals, and sometimes a competitor you had never heard of dominates the answers.
Record all four with a date attached. A baseline without a date is not a baseline.
Choosing prompts you can defend
The prompt set is the measurement instrument, so a biased set produces a flattering, useless number. Four rules keep it honest.
Write the questions a stranger would ask. Not "is [your brand] good" — nobody undecided types that. "Best [category] for [situation]", "who should I hire to [job]", "is [category] worth the cost" are the real shapes.
Cover the intent spread. Roughly: a third informational ("how does X work"), a third comparison ("best X for Y", "X vs Y"), a third decision ("who should I hire in [place]", "is X worth it"). Vendors compete well in the first third and poorly in the last, so a set weighted to informational prompts will overstate your position considerably.
Include prompts you expect to lose. A set you win is a set you cannot learn from. The losses are where the diagnostic information lives.
Freeze it. Changing prompts between runs destroys comparability — the fastest way to produce a chart that means nothing while looking like progress. Version the set, keep a stable core across runs, and add new prompts as a clearly separated cohort rather than by editing the originals.
Our companion piece on building a prompt set for AI visibility tracking covers sourcing and phrasing the questions themselves; this is about the statistical hygiene around them once chosen.
Reading movement through the noise
Non-determinism is the practical obstacle to trusting anything. Four habits handle it.
| Problem | Bad practice | Better practice |
|---|---|---|
| Same prompt, different sources each run | Report the latest run as the result | Report a rolling average across three or more runs |
| One engine moves, others do not | Average everything into one score | Chart per engine; investigate the mover |
| Small prompt set, jumpy percentages | Read every wobble as a trend | Use enough prompts that one flip is not a visible spike |
| Prompt set edited mid-programme | Compare the new number to the old | Version the set; annotate the chart where it changed |
A practical threshold: with a frozen set of 15 or more prompts run weekly, a change is worth investigating when it holds for three consecutive runs. Anything shorter is inside the noise, and treating it as signal produces a lot of confident, wrong explanations.
The mirror-image error is dismissing everything as noise. A prompt that flips from never-cited to cited in five of six consecutive runs has genuinely changed, and it is usually worth knowing exactly which page the engine started pulling from.
What counts as a good number in practice
With the caveat that these are orientation rather than published research, here is how to read your own figures.
Citation rate under 5% across a well-constructed set usually means a structural problem rather than a content one — most often that the retrieval crawlers cannot reach you, or your content does not render without JavaScript. Check crawler access and rendering before writing anything new. Near-zero is a symptom, not a starting point.
5% to 20% is where most sites sit once access is clean and some content exists. You are a plausible candidate the engines select occasionally. This is where depth and corroboration move the number.
Above 30% on a genuinely mixed set means you are a default answer in your category. The work shifts from earning citations to holding them — monitoring for regressions, keeping facts current, and watching the competitor set for new entrants.
Share of voice is the more stable frame. If three named competitors split most of the mentions and you are fourth with 9%, that is a specific, actionable gap. If you are at 9% and the leader is at 11%, your market is fragmented and the opportunity is much larger than the absolute number suggests.
Benchmarking against competitors, properly
Do not simply track your own score. Run the same frozen prompt set and record every brand named, then track the whole leaderboard over time.
This gives you three things one-brand tracking cannot.
Context for your own movement. A citation rate that fell while every competitor's also fell is an engine behaviour change, not your failure. A rate that held while a rival's doubled is a competitive loss disguised as stability.
An early warning on new entrants. A brand appearing in your answer set for the first time and climbing is worth investigating immediately, while the pattern is small enough to learn from cheaply.
A concrete target. "Reach the citation rate the current second-place brand holds" is a goal with a number attached that a team can actually work toward, unlike "improve AI visibility".
If a competitor is consistently ahead, the diagnostic questions — what does the engine know about them that it does not know about you — are laid out in when a competitor is cited in AI answers and you are not. Very often the answer is off-site: the corroborating sources covered in off-site GEO.
Turning a baseline into a programme
A baseline is only useful if the same measurement repeats. That means fixing four things and writing them down: the prompt set version, the engine list, the cadence, and the definition of a citation (named anywhere in the answer, versus linked as a source — both are valid, but pick one and stay with it).
Then the loop: measure, identify the prompts you lose, fix the specific reason you lost them, re-measure. The fixes are ordinary — access, rendering, a page that answers that exact question directly, better third-party corroboration — but doing them against measured losses rather than intuition is what separates a programme from a hobby.
DigiRank's AI Visibility Tracker is built around this loop: a frozen prompt set re-run on a schedule across all six engines, per-engine citation rate and share of voice charted over time, the full competitor set extracted from the answers rather than supplied by you, and an Opportunities view that lists the prompts you are missing from with a one-click route to a drafted page. Tracking starts on the $99/mo Starter plan with 200 prompt-checks a month — enough for a 15-prompt set across six engines run fortnightly — and Agency at $249/mo raises that to 2,500 with white-label share pages for clients. What to demand from any tool in this category, including this one, is in the AI visibility tracker buyer's guide, and the mechanics of what the tracker does are in how an AI search visibility tracker works.
Frequently asked questions
What is a good AI citation rate? There is no published industry benchmark, and any figure quoted as one is invented. Judge your rate against the competitors the engines actually name for your prompt set. As orientation: under 5% on a well-built set usually indicates a structural problem such as blocked crawlers rather than a content problem; 5–20% is common once access is clean; above 30% on a mixed set means you are a default answer in your category.
How do I calculate share of voice in AI search? Count every brand mention across the answers to your frozen prompt set, then take your mentions as a proportion of that total. Unlike a raw citation rate it moves relative to competitors, so it stays meaningful when engine behaviour or market conditions shift.
How many prompts do I need for a reliable baseline? Fifteen or more across a mix of informational, comparison and decision intents. Below about ten, a single flip produces a visible percentage swing and you will spend your time explaining noise.
How do I choose which prompts to track? Write the questions a stranger would actually type before choosing a supplier, spread them across informational, comparison and decision intent, and deliberately include prompts you expect to lose — those carry the diagnostic value. Never include prompts containing your own brand name in a baseline set; they inflate the score without informing anything.
How often should I re-run the set? Weekly during active work, monthly for maintenance. Weekly gives enough runs to distinguish a real trend from non-determinism inside about six weeks.
Why does my score change when I haven't changed anything? Because AI answers are non-deterministic — the same prompt can surface different sources on different runs. Treat any single run as a sample, use a rolling average across three or more, and only investigate a change that persists for three consecutive runs.
Should I average my score across engines? No. Chart each engine separately. Averaging hides the most useful event in the data: one engine moving while the others hold, which usually points at a specific, fixable cause.
Can I change my prompt set later? Yes, but treat it as a version change rather than an edit. Keep a stable core for comparability, add new prompts as a separate cohort, and annotate the chart at the point the set changed so nobody reads a methodology change as a result.
