Measurement

Why AI Answers Change Between Runs, and How to Track Them Anyway

A single AI answer is a sample, not a ranking. Treat it like one and most of the confusion in visibility tracking disappears.

By DigiRank Expert · October 2, 2026

Several glass marbles of different colours scattered across a sheet of graph paper on a desk

Short answer, as of October 2026: AI assistants are built to vary their output, so the same prompt asked twice can name different businesses — which means a single answer tells you almost nothing, and visibility has to be measured as a rate across repeated runs rather than as a position. Once you treat each answer as one sample from a distribution, the readings stop looking erratic and start looking like data.

This is the part of AI visibility tracking that most surprises people who come from rank tracking. A search result page is close to deterministic: ask twice, get nearly the same ten links. An AI answer is generated fresh each time. If your reporting assumes the first and you are measuring the second, every week will look like a crisis or a triumph, and neither will be real.

The four sources of variance

It helps to separate them, because they call for different responses.

1. Sampling in the model itself

Language models produce text by choosing each next word from a range of plausible options, with some randomness deliberately included. That randomness is why the wording differs between runs, and it is also why a list of "three good options" may contain different names. When several businesses are roughly equally supported by the available evidence, which ones make the list is partly chance.

This is the variance you cannot remove. You can only average over it.

2. Retrieval differences

Assistants that search the web before answering do not always retrieve the same pages. The search step can be phrased differently from run to run, the underlying index changes, and the set of pages passed to the model shifts. If your page is in the retrieved set you may be cited; if it narrowly misses, you will not be, regardless of how good the page is. How that selection works is covered in how ChatGPT, Perplexity and Gemini pick their sources.

Retrieval variance is the kind you can influence, because being retrieved consistently is largely a matter of having the clearest page for the question.

3. Context the assistant has about the asker

Location, conversation history, account settings and whether the user is signed in can all change an answer. A prompt asked from one city may surface different local businesses than the same prompt asked from another. This is not noise in the statistical sense — it is a real difference between audiences — but it looks like noise if your tracking does not hold these conditions constant.

4. Changes to the product

Assistants are updated often, sometimes without announcement. A model version change or an adjustment to how search is used can move results for everyone at once. This is the easiest kind to recognise: your competitors move at the same time you do.

What variance looks like in practice

A worked illustration, with invented numbers for clarity rather than measured ones. Suppose you ask one prompt ten times in one assistant on the same day:

Outcome across 10 runsWhat a single run would have told youWhat it actually means
Named in 9 of 10"We are cited" (probably)A stable position; you own this answer
Named in 6 of 10"We are cited" or "we are not", close to a coin flipContested; you are one of several interchangeable options
Named in 2 of 10Most likely "we are not cited"Marginal; occasionally retrieved, rarely chosen
Named in 0 of 10"We are not cited"Absent; a retrieval or relevance problem

The second and third rows are where single-run tracking misleads. A business named six times in ten will, on any given weekly check, appear to flip between cited and not cited. A chart of those single checks looks like volatility. The underlying position has not moved at all.

The useful number is the mention rate: the share of runs in which you were named. It is the AI-search equivalent of a position, and unlike a single reading it is stable enough to compare month to month.

How many runs are enough

There is no universal answer, but there is a practical one.

One run per prompt is not a measurement. It is an anecdote. Do not report it as a trend.

Three to five runs per prompt per period is enough to sort prompts into the rough bands in the table above — owned, contested, marginal, absent — which is the level of precision most decisions need. You do not need to know whether your mention rate is 60 or 65 per cent. You need to know that the prompt is contested rather than owned.

More runs matter only when you are measuring a change. If you rewrote a page and want to know whether it worked, the before and after readings each need enough runs that the difference is larger than the wobble. A move from two-in-five to three-in-five is not evidence of anything. A move from one-in-five to five-in-five is.

The cost of repetition is real, because trackers meter on prompt-checks — one prompt, one engine, one time. Twenty prompts across six engines with three runs each is 360 checks per period. That arithmetic is the honest reason to keep a prompt set small and well chosen rather than large and thinly sampled: fifteen prompts measured properly beat sixty measured once. Building a set that size is the subject of how to build a prompt set.

Aggregate before you conclude

Individual prompts are noisy. Groups of prompts are much less so. Three habits take most of the noise out of a report.

Report by prompt group, not by prompt. Combine all your cost prompts, all your comparison prompts, all your problem prompts. A group-level mention rate built from eight prompts and several runs each is steady enough to trend.

Compare periods, not days. Month against month, using all runs within each. Weekly lines are for spotting breakage, not for judging progress.

Always include a competitor. If your rate and a competitor's rate fall together, the assistant changed. If yours falls while theirs holds or rises, something about your presence changed. Without the comparison line you cannot tell these apart, and they call for opposite responses. This is the same logic as a share-of-voice baseline, set out in AI visibility benchmarks.

Telling a real change from noise

A short diagnostic for the Monday-morning question, "we dropped — is it real?"

Did it persist? A drop that is still there after two more measurement periods is a change. One that reverses is variance.

Is it concentrated? Real changes usually hit a cluster of related prompts — every prompt served by one page, or every prompt in one topic. Noise is scattered evenly.

Did the cited sources change? This is the strongest signal. If the prompts where you disappeared now consistently cite a new source in your place, something displaced you. If the sources are the same mix as before and you simply were not in this sample, wait.

Did everyone move? Check the competitor line. A simultaneous shift across brands points to the assistant, not to you.

Did something change on your side? A redesign, a migration, a changed robots rule, a page whose opening paragraph was edited. Access regressions produce sudden, total drops rather than gradual ones, and are worth ruling out first.

If the answer is "persisted, concentrated, new sources" you have a genuine loss, and the next step is the diagnostic sequence in citation decay. If it is "reversed, scattered, same sources" you have a normal week.

What to tell a client or a manager

Variance is uncomfortable to explain to someone who expects a ranking. Three sentences usually do it.

First: an AI answer is generated fresh each time, so we measure how often we appear rather than where we appear. Second: we report that rate by topic and by month, because individual answers bounce. Third: we watch the sources cited alongside us, because that is what tells us whether a movement is real.

Then show bands rather than decimals. "We own four of our eight cost prompts, contest three and are absent from one" is a statement a non-specialist can hold on to. "Visibility moved from 41.3 to 38.7" invites a conversation about a difference that is almost certainly inside the noise.

Where a tracker helps

Doing repeated runs by hand across several engines stops being practical almost immediately, which is the legitimate case for software. DigiRank Expert's AI Visibility Tracker runs your prompt set on a schedule across ChatGPT, Claude, Gemini, Perplexity, Grok and Google AI Mode, stores the answer behind each reading, and charts visibility over time from a daily snapshot job so that you are looking at a trend rather than a single draw. Tracking starts on the $99/mo Starter plan with 200 prompt-checks a month; the Agency plan raises that to 2,500, which is the allowance at which repeated runs across a full prompt set become comfortable. The rest of the modules are listed on the features page.

Whatever tool you use, ask it one question: how many times is each prompt run before a number is shown? If the answer is once, read every chart it produces as a set of anecdotes.

Frequently asked questions

Why does ChatGPT give a different answer every time I ask the same question? Because the model selects its wording with deliberate randomness, and when it searches the web it may retrieve a different set of pages each time. When several businesses are about equally well supported, which ones are named is partly chance. This is by design and affects every assistant.

How do you measure AI visibility if the answers keep changing? As a mention rate: the share of repeated runs in which your brand is named for a prompt or a group of prompts. A rate across several runs is stable enough to compare month to month, whereas a single answer is not.

How many times should each prompt be run? Three to five runs per prompt per period is enough to tell whether you own, contest or are absent from an answer. More runs are needed only when you are measuring the effect of a specific change and need the difference to exceed normal variance.

How can I tell if a drop in AI visibility is real? Check whether it persists across further measurement periods, whether it is concentrated in related prompts, whether a new source is now cited in your place, and whether competitors moved at the same time. A persistent, concentrated drop with new sources is real. A scattered drop that reverses is variance.

Does my location change the AI answer I get? It can. Assistants may use location, sign-in state and conversation history, so the same prompt can surface different businesses for different users. Consistent tracking holds those conditions constant so that the comparison over time is fair.

Is a visibility score from a single run useful? Only as a spot check that something is not broken. A single run cannot distinguish a brand that is named nine times in ten from one named four times in ten, so it should never be reported as a trend.

Why did all my prompts change on the same day? A simultaneous shift across many prompts, especially one that also affects competitors, usually means the assistant itself was updated. Compare against a competitor line before concluding that anything changed on your side.

See where you stand across 6 AI engines.

DigiRank tracks whether ChatGPT, Perplexity, Gemini, Copilot, Claude, and Grok cite you — then ships the Princeton-scored content that wins the citation.

Start 14-day free trial