How consistent are AI recommendations, really?
Last updated: September 14, 2026
Why this matters if you want to be found by AI
The AI shortlist is not a ranking. It is a re-roll. People now ask an assistant which tool or company to pick and take the names it hands them, so being on that list is real distribution. But if the list changes every time it is generated, "ChatGPT recommended us once" means almost nothing. What matters is whether you land in the part that stays put. That is the whole question, and it has a measurable answer.
What we did
Ten open "best X for a small business" questions (PR distribution, CRM, email marketing, accounting, e-commerce and more), three times each, to two web-search-enabled assistants: ChatGPT via the OpenAI Responses API with web search, and Gemini with Google Search grounding. Sixty responses, one day. From each answer we extracted the businesses it recommended and measured how much that set overlapped across the three identical runs (mean pairwise Jaccard overlap). All 290 extracted names appear verbatim in their source answers. Method, code and every response are public.
What we found
| Assistant | Consistency | Range | Searched the web |
|---|---|---|---|
| ChatGPT (web search) | 87% | 59–100% | 100% |
| Gemini (Search grounding) | 52% | 33–67% | 97% |
Overall consistency was 69.5%. Ask again minutes later and roughly 31% of the set is different. The useful part is the shape. Of 127 distinct business slots, 56% were named in every run and 27% appeared in only one of three. The first one to three recommendations recurred every time; the fourth, fifth and sixth rotated. The names that recurred most were the big, heavily-described brands in each category: Shopify was named 12 times, Wix 8, Squarespace 7.
The stable core is not random. It is the most-mentioned brands.
That last detail is the finding people skip past, so here is the reasoning laid out.
How we got here
Our data: the brands that survived all three runs were the ones the web describes most often (Shopify, Wix, Squarespace and their equivalents in each category). The churning tail was thinner-footprint names. Sixty responses, September 2026.
Ahrefs, December 2025, 75,000 brands: branded web mentions are the strongest non-video predictor of AI visibility, at a Spearman correlation of 0.664 on ChatGPT, 0.709 on AI Mode, 0.656 on AI Overviews. And the same brands show up across platforms: visibility on AI Overviews and AI Mode correlates 0.821, AI Mode and ChatGPT 0.769.
Put together: our "locked core" and Ahrefs' "consistently visible across platforms" are the same phenomenon seen twice. A brand that is described often, near its topic, on many independent sites gets named every time and on every engine. A brand with a thin mention footprint gets named sometimes, which is what a churning tail looks like from the inside. Neither dataset says this on its own. Ahrefs measures whether you appear; ours measures whether you keep appearing. The overlap is the claim: mentions do not just get you into the shortlist. They decide whether you are in it every time.
One caveat, once: our study is ten questions on a single day and Ahrefs measures visibility rather than run-to-run stability, so read this as strongly consistent with, not proven by. The direction we would stake money on. The exact percentages we would not.
Do the engines agree with each other?
Mostly not, and it shows up in both datasets. Our Gemini runs were far less stable than ChatGPT's (52% against 87%) and frequently returned different businesses for the same question. Ahrefs sees the same factors matter on every platform but at different strengths (mentions correlate 0.709 on AI Mode, 0.664 on ChatGPT, 0.656 on AI Overviews). Treating "AI visibility" as a single number hides this. There is no unified rank; there is per-engine, per-run behavior, and the same brand can be locked on one engine and rotating on another.
One honest disagreement
Does a third of the answer re-rolling mean AI recommendations are too unreliable to bother optimizing for? The camps split:
School A: it is noise, wait it out
If a third re-rolls every time and engines change behavior on each model update, chasing a seat is chasing a moving target. Spend on channels you can measure and hold.
School B: the stable core is the whole point
The churn is in the tail; the head is locked. That stability at the top is exactly what makes the position worth earning, and it is bought with mentions, which you can build.
School B is right, and School A is quietly moving the goalposts. "It's noise" only holds if the whole answer re-rolls. It doesn't. The head is locked, the tail churns, and the head is made of the most-described brands. That is not a coin flip. That is a lever.
Related entries
Changelog
September 14, 2026: Rewritten to add the cross-read against Ahrefs' 75,000-brand study (the "stable core is the most-mentioned brands" finding). First published earlier the same day.
Sources (primary): Pressfront Research, Run-to-Run Consistency of Business Recommendations from Web-Search-Enabled Large Language Models, 60 responses, September 13, 2026, DOI 10.5281/zenodo.22738861, data and code at GitHub; Ahrefs, "Top Brand Visibility Factors in ChatGPT, AI Mode, and AI Overviews", 75,000 brands, Spearman correlations, December 12, 2025.