Pressfront Research

How consistent are AI recommendations? A first look.

Published September 13, 2026 · An exploratory study by Pressfront

We asked ChatGPT and Gemini the same ten "best X for a small business" questions, three times each, within a single day. About a third of the businesses they recommended changed between identical runs. The most-named businesses stayed locked in every time; the rest of the shortlist shuffled underneath them.

~31%
of recommended businesses changed when the same question was asked again
56%
of businesses named were named in every run; a "locked" core
27%
were named in only one of the three runs; a shuffling tail

Why we ran this

People increasingly ask an AI, rather than a search engine, which company or product to use. So a fair question for any business owner is: if a customer asks ChatGPT or Gemini to name the best option in my category, how stable is that answer? If the same question produces a different shortlist each time, then "being recommended by AI" is less a fixed ranking and more a moving target. We wanted a number for it, so we measured it.

What we did

We chose ten open buying questions across ten categories (PR distribution, CRM, email marketing, project management, accounting, marketing agencies, website builders, live chat, e-commerce and scheduling), phrased the way a small-business owner would type them. We deliberately used open "best X" questions, not "X vs Y" comparisons, because a comparison can only return the two names in the question. We asked each question three separate times on both ChatGPT (with web search) and Google Gemini (with Search grounding), all within one day. For each answer we extracted the businesses it recommended, then measured how much the recommended set overlapped across the three identical runs. Full method and limitations are at the bottom.

The headline number

Averaged across every question and both engines, only about 69% of the recommended businesses stayed the same from one run to the next. Put the other way: roughly one in three names changed when we asked the exact same question again, minutes apart, with nothing changed but the click.

The leaders lock in. The tail shuffles.

The churn was not evenly spread. When we looked at every business named for a given question and counted how often it reappeared across the three runs, a clear pattern showed up: 56% of the businesses were named in all three runs, a stable core the model keeps returning to. But 27% were named in only one of the three runs, a rotating tail that changes each time you ask. The top two or three recommendations were remarkably steady; the fourth, fifth and sixth slots were where the shuffle happened, and that is exactly the space a smaller or newer business competes for.

ChatGPT was steadier than Gemini

Both engines searched the web on every run and nearly always produced a real shortlist, so this is a like-for-like comparison of what they recommend.

EngineAvg. consistencyGave a shortlist
ChatGPT (web search)87%100% of runs
Gemini (Search grounding)52%97% of runs

ChatGPT tended to repeat itself closely, its top recommendations were near-identical run to run (per-question consistency ranged from 59% to 100%). Gemini kept its leading picks but reshuffled the rest of the list more freely (33% to 67%), which is why its overlap score is lower. Neither is "wrong"; they simply treat the tail of a recommendation list as more negotiable than the head. With only ten questions per engine these averages carry no formal margin of error, so treat the gap as a clear direction rather than a precise figure.

Consistency per question

QuestionChatGPTGemini
Best PR distribution services59%50%
Top CRM platforms for startups100%42%
Best email marketing services62%63%
Leading project management software100%67%
Best accounting software for freelancers100%33%
Best digital marketing agencies89%61%
Top website builders100%62%
Best live chat software100%58%
Top e-commerce platforms100%33%*
Best scheduling software62%50%
Consistency = share of recommended businesses that overlapped across three identical runs. With three runs, per-question values fall on a coarse scale, so read these as directional, not precise. *The e-commerce cell for Gemini is based on two runs, not three: one of its three answers named only a single business and was excluded by the two-or-more rule.

The businesses named most often

Across all runs and both engines, and after excluding any name we had put in a question ourselves, the three most-named businesses were Shopify (12 mentions), Wix (8) and Squarespace (7). Behind them, twelve businesses tied at six mentions each: EIN Presswire, Pipedrive, Attio, Klaviyo, Trello, ClickUp, Monday.com, FreshBooks, Hostinger, Tidio, Tawk.to and Acuity Scheduling. These formed the stable core, recommended by both engines across most runs without prompting. The instability lived below them.

What this means if you run a business

Being named by AI is not a ranking you win once. On open "best X" questions the shortlist reshuffles, and the churn is worst in exactly the positions a challenger can realistically reach. The businesses that appear every single time are the ones described, by name, across many credible independent sources, so the model keeps arriving at the same answer. That consistency is the thing you actually build toward.

See where you stand, freeA live AI Visibility Report across ChatGPT, Gemini and Google AI. About a minute.

Method & limitations

Cite as: Pressfront Research, "How Consistent Are AI Recommendations? A First Look", September 2026, pressfront.co/ai-visibility-index.html

Related: What is GEO? · Is my business on ChatGPT? · Why doesn't AI recommend my business?

Pressfront LLC · Tampa, FL · pressfront.co