Why testing ten prompts tells you nothing
Most brands check their AI visibility by asking an assistant a handful of questions and reading the answers. It feels like evidence. It is closer to taking one sample from a distribution nobody has characterised.
Start with the uncomfortable part: identical inputs do not give identical outputs, even when everything is configured for determinism. A 2024 study ran five models across eight tasks, ten times each, at temperature zero and other determinism settings. Accuracy varied by up to 15% across runs, the gap between best and worst run reached 70% on some tasks, and “none of the LLMs consistently delivers repeatable accuracy across all tasks” (Atil et al., arXiv:2408.04667). Temperature zero does not buy you a stable answer.
So one run is a coin flip. The obvious fix is to run the same prompt many times. That turns out to be close to the least useful thing you can do.
Where the variance actually lives
A 2026 preprint decomposed the variance in brand mentions across 12,933 responses, 20 brands, eight languages and three models. Resampling the same prompt accounted for 34.8% of the variance in a single response, and the query’s language for 26.5%. Brand identity, the thing you are trying to measure, accounted for 1.5%. Repeating an identical prompt one more time reduced relative error variance by 0.0003 (Żatuchin, arXiv:2607.13304).
That is a single-author preprint on one dataset, so hold the exact percentages loosely. The qualitative conclusion is what matters, and it is hard to argue with: reliability comes from spreading your sampling across phrasings, languages and engines, not from repeating one question.
What that implies for a measurement design
- Sample the phrasing space, not the prompt. Your customers do not share your vocabulary. “Best luxury watch for resale value” and “which watches hold their value” are the same intent and can return different brands.
- Treat each engine separately. Overlap between engines is partial at best, as the work on citation sources shows. Averaging them hides the engine where you are absent.
- Sample languages if you sell in them. Language was the second largest variance component in that decomposition, ahead of everything about the brand.
- Report intervals, not points. “We appear in 22% of answers, plus or minus 6” is a finding. “We appear in ChatGPT” is an anecdote.
- Keep the sample fixed over time. If the prompt set changes with every report, you cannot tell a real movement from a new question.
None of this requires exotic tooling. It requires accepting that the unit of measurement is a distribution over many phrasings, and that a number without a sample size attached is decoration.
Once you have that, the question becomes which numbers to compute, and there are three worth having before any others. It also becomes possible to tell whether a change you made actually moved anything, which single spot checks can never do.