Our research

How often does the "best X" recommendation change when you just ask again?

Last updated: September 14, 2026

About a third of it changes. We asked ChatGPT and Gemini the same open "best X for a small business" questions three times each. Roughly 31% of the recommended businesses differed between identical runs (69.5% overlap). ChatGPT held 87% run-to-run consistency; Gemini 52%. So here is the consequence, stated plainly: one AI check is a sample of one from a process that re-rolls. "ChatGPT recommended us" from a single ask, or "we didn't show up" from a single ask, is not a measurement. Ask at least three times before you believe either direction.
~31%of the recommended set changed on an identical repeat (60 responses)
1 of 3the sampling floor: below this you are reading noise
27%of all named slots appeared in only one of three runs (pure tail)

Why one check lies to you in both directions

People treat a single AI answer like a scoreboard. It isn't one. It is a draw. Ask "best CRM for a small business," get a list; ask again a minute later, get a list that is roughly 31% different. That cuts two ways, and both are traps.

If you show up once, it is tempting to screenshot it and call it proof. But a name that appears in only one of three runs is, by our count, the single most common kind of name in the whole dataset: 27% of all recommended slots appeared exactly once. A one-time appearance is the statistical signature of the churning tail, not of a won position. And if you don't show up once, that is not proof you're invisible either. You might be a name that surfaces in two runs out of three and you happened to catch the third. Either way, n=1 tells you almost nothing.

What we measured

Ten open "best X for a small business" questions (PR distribution, CRM, email marketing, accounting, e-commerce and more), three times each, to two web-search-enabled assistants: ChatGPT via the OpenAI Responses API with web search, and Gemini with Google Search grounding. Sixty responses, one day. From each answer we pulled the businesses it recommended and measured how much that set overlapped across the three identical runs (mean pairwise Jaccard overlap). All 290 extracted names appear verbatim in their source answers. Everything is public.

AssistantRun-to-run consistencyRange across questions
ChatGPT (web search)87%59–100%
Gemini (Search grounding)52%33–67%

Overall overlap was 69.5%, so ~31% of the set churns on a repeat. And the churn isn't evenly spread. On Gemini the answer was less than half stable on some questions (33% overlap on one). If your one check happened to land on Gemini, your n=1 is even shakier than the average suggests.

So how many checks is enough?

Three is the honest floor for a category, not because three is magic but because it is the smallest number that lets you tell a locked name from a lucky one. A name in all three runs is a real position. A name in one of three is probably noise. A name in two is a maybe you should keep watching. The point is not the exact count. The point is that you cannot get any of that from a single ask, and most people optimizing for AI visibility are staring at a single ask.

How we got here

Our data: across 127 distinct business slots, 56% were named in every run (locked) and 27% in only one of three (pure tail). Overall run-to-run overlap 69.5%, so about 31% of the set re-rolls. Sixty responses, September 2026, DOI-registered.

The reasoning: if the most common outcome for an individual name is a single appearance across three tries, then any conclusion drawn from a single try is drawn from the noisiest possible slice of the data. You cannot distinguish "we won a seat" from "we caught a re-roll" without repeating the draw. That is not a modeling assumption. It falls straight out of the churn number.

One caveat, once: this is ten questions on a single day, sixty responses. Directional, not definitive. The specific 31% will move with the sample. The rule it implies, do not trust one check, does not.

Want a real read instead of a single lucky (or unlucky) screenshot? Pressfront's free AI Visibility Report runs the actual engines against your category so you see whether you're in the locked core, the churning tail, or missing.

Related entries

Changelog

September 14, 2026: First published. Isolates the churn figure from the broader consistency study and draws out the single practical rule (never trust one check).

Source (primary): Pressfront Research, Run-to-Run Consistency of Business Recommendations from Web-Search-Enabled Large Language Models, 60 responses, September 13, 2026, DOI 10.5281/zenodo.22738861, data and code at GitHub.