Measurement

How do I actually measure whether AI is recommending my brand?

Last updated: September 14, 2026

Not with a single check, because a single check lies. When we asked ChatGPT and Gemini the same "best X" questions three times each, about 31% of the recommended businesses changed between identical runs (69.5% overlap; ChatGPT 87% consistent, Gemini 52%). So a screenshot proves nothing. The method that works: write the exact prompts a customer would type, ask them cold in a fresh session with no memory or personalization, repeat each prompt several times because the answer re-rolls, run it on every engine your buyers use because they diverge, and log for each run whether you appeared, in what slot, and who got named instead. Your visibility is a rate across samples, not a yes or no from one answer. You do not need a paid dashboard for this. You need a spreadsheet and the discipline to run it the same way every time.
~31%of recommended businesses changed on an identical repeat ask (our data, 60 responses)
87% / 52%ChatGPT vs Gemini run-to-run consistency: the engines are not the same instrument
N ≥ 3minimum repeats per prompt before a single answer means anything

Why one check is worse than no check

One check gives you a number you will trust and it is often wrong. Ask ChatGPT "best CRM for a small business," see your name, feel good, and close the tab. Ask again ten minutes later and there is roughly a one in three chance the set of names moved. If your brand was in the churning part of the answer, you just recorded a win that will not repeat. If you were absent on the one run you happened to look at but present on the next two, you recorded a loss that is not real. A measurement you cannot reproduce is not a measurement. It is a mood.

This is the thing the "just ask ChatGPT" advice and half the paid dashboards quietly skip. Non-determinism is not an edge case here. It is the center of the problem, and it dictates the whole method.

Why you sample across runs AND across engines

Two separate sources of variation, and you have to control both.

How we got here

Variation within an engine (run to run): our own study, 60 responses over one day, found 69.5% overlap on identical repeats, meaning about 31% of the answer re-rolls. That alone kills the single check: you need several samples of the same prompt to separate the stable core from the churning tail.

Variation between engines: in the same study ChatGPT held 87% consistency and Gemini only 52%, and the two frequently returned different businesses for the same question. Ahrefs, across 75,000 brands, sees the same signals matter on every platform but at different strengths per engine (mentions correlate 0.709 on AI Mode, 0.664 on ChatGPT, 0.656 on AI Overviews). You can be locked on one engine and rotating on another.

Put together: a real measurement has to sample twice over. Repeat each prompt to average out run-to-run churn, and run every engine to catch per-engine divergence. Collapsing "AI visibility" into one number from one engine on one run hides both sources of error at once. The method is not "check if AI mentions you." It is "estimate the rate at which each engine mentions you, with enough samples that the rate is stable."

The actual method

You can run this by hand today. No tool required.

  1. Write the prompts your buyers type. Not "is Pressfront good." The open questions a customer asks before they know you exist: "best press release distribution for a startup," "how do I get my company recommended by AI," "cheapest way to get media coverage." Ten to twenty of them. These are your test set, and you freeze it so runs are comparable over time.
  2. Ask cold. Fresh session, logged out where you can, no custom instructions, no memory, no prior chat in the thread. If the assistant knows you or you primed it, you are measuring your own footprint, not the market's. Turn personalization off.
  3. Repeat each prompt at least three times. Three is the floor; five is better. New session each time. This is the step everyone skips and it is the one that makes the number mean something.
  4. Run every engine your buyers actually use. ChatGPT, Gemini, Perplexity, Copilot, whatever your audience opens. They diverge, so one engine is one data point, not the answer.
  5. Log four things per run: did your brand appear (yes/no), in what position, which competitors were named, and whether the assistant searched the web or answered from memory. A row per prompt-per-engine-per-run. A spreadsheet is enough.
  6. Report a rate, not a verdict. "Named in 7 of 15 ChatGPT runs, position 4 to 6, Shopify named every time" is a measurement. "ChatGPT recommends us" is a screenshot. The rate is what you track month over month to see if your mention-building is working.

What a paid dashboard does and does not buy you

Tools that track AI visibility mostly automate exactly the loop above at scale: more prompts, more engines, more repeats, charts over time. That is genuinely useful once you are running hundreds of prompts and cannot do it by hand. What a dashboard does not do is change the underlying method, and a tool that reports a single visibility score without telling you how many samples it ran per engine is hiding the one number that decides whether the score is trustworthy. Ask any vendor how many repeats per prompt sit behind their figure. If the answer is one, the figure is a mood with a logo on it.

Want the loop run for you against the live engines, with the sampling built in? Pressfront's free AI Visibility Report checks the real assistants for your category and shows you where you land, and which competitors get named instead.

Related entries

Changelog

September 14, 2026: First published. Method grounded in our run-to-run consistency data (the case for sampling) cross-read with Ahrefs' per-engine correlations (the case for running every engine).

Sources (primary): Pressfront Research, Run-to-Run Consistency of Business Recommendations from Web-Search-Enabled Large Language Models, 60 responses, September 13, 2026, DOI 10.5281/zenodo.22738861, data and code at GitHub; Ahrefs, "Top Brand Visibility Factors in ChatGPT, AI Mode, and AI Overviews", 75,000 brands, Spearman correlations, December 12, 2025. AI behavior evolves, so this entry is dated and revisited.