From Mode’s research series. This one is built on our own measurement equipment rather than somebody else’s statistics, and it includes a result that is unflattering to us. The method is written out at the bottom so you can check it or repeat it.
Somebody sends you a report. It says ChatGPT does not mention your business. Or, better, that it does. Either way there is a question worth asking before you act on it, and almost nobody asks it: how many times did they look?
We have been in a position to answer that properly since the middle of July, because we built a machine that asks the same questions over and over. Here is what 17 days of it found.
What we did
Mode runs a scheduled harness that puts a fixed bank of 53 questions to the AI engines through their official APIs and records every URL each answer cites. Some questions run daily, the rest run three times a week. Between 13 and 30 July 2026 it made 723 calls, of which 581 came back clean. That gave us 56 question and engine combinations measured at least five separate times, and 380 pairs of back to back measurements to compare. In total the engines cited 1,481 different websites.
Then we asked the simplest possible question of that data. Ask the same engine the same thing again a day or two later. Does it point at the same sources?
It mostly does not
Across those 380 consecutive pairs, on average only 61% of the sources named across any two runs appeared in both of them. Put the other way round, roughly four in ten of the sources involved turned over between one measurement and the next. Only 4.5% of consecutive measurements returned an identical list. Seventeen pairs out of 380.
Some of that is an artefact worth owning. The number of sources an answer cites is itself variable, averaging 12.8 and ranging from 3 to 22, and only a third of consecutive pairs cited the same number of sources at all. A run that names 8 sites and a run that names 16 will look unstable even if the smaller list sits entirely inside the larger. So we recalculated using containment instead: of the sources in whichever run named fewer, how many also appeared in the other? That comes out at 77.4%. Even after controlling for the count, about 23% of the shorter list’s sources are not shared. The churn is smaller than the headline suggests, and it is still real.
The longer view is the one that would worry me if I were buying. Of those 1,481 domains, only 20.7% appeared on every single measurement of their question. 26.7% appeared exactly once and were never seen again. There is a stable core of sources the engines keep returning to, and around it a large rotating cast of sites that surface once and vanish. Of the sources cited on the very first measurement of each question, 62% were still being cited on the last one 17 days later. So roughly a third of what was true about a question in mid July was no longer true by the end of the month.
The part that matters commercially
We group the question bank by type, and the differences between groups are sharper than the average lets on.
The least stable group, by a distance, is the commercial one: the who-should-I-hire questions. Those turned over 58% of their sources between consecutive measurements. Questions about verticals came next at 48%, then trust questions at 42%. The most stable group was plain cost questions at 32%.
Read that ordering again, because it is the finding. The questions where an AI is effectively making a recommendation, the ones that decide who gets the enquiry, are the ones whose sources move around the most. A question like whether there is a UK agency that builds websites in five days pulled 50 different websites across seven measurements, with barely three in ten sources holding from one run to the next. The further a question moves from a fact and towards a judgement about who to use, the less settled the machine is about where to look.
Our own result, including the bad half
We measure ourselves with the same equipment, so here is our own scorecard. Across 581 clean runs, Mode Marketing was named in the text of 148 answers. A Mode web page was cited as a source in 8. On Perplexity we were named 98 times and cited zero times. On Claude, 30 times and zero. All eight citations came from one engine.
That gap is worth understanding, because a lot of AI visibility marketing quietly blurs it. Being talked about is not the same as being linked to. An engine can describe your business accurately, from memory, and send the reader nowhere near your website. We would rather publish that than a flattering number.
What a business owner should take from this
Three practical things.
Treat any one AI visibility check, including one of ours, as a sample rather than a verdict. If a report tells you an AI does not mention you, that is one roll of a fairly loose die. If it tells you it does, the same applies. The honest unit of measurement here is a trend across repeated runs, not a screenshot.
Ask any provider how many times they measured and across what period. It is a fair question and the answer is revealing. If a tool or an agency is producing a confident visibility score off a single pass, the score is carrying more precision than the underlying data can support.
And do not panic at movement, or celebrate it. Appearing once, or disappearing once, is inside the normal range of noise we measured. What is worth acting on is a direction that holds over weeks. The underlying work that earns a place in the stable core, being readable by machines, publishing verifiable facts, being genuinely worth citing and being mentioned elsewhere on the open web, is slow and unglamorous and does not respond to a single week’s reading either.
Where Mode fits, stated plainly, including against our own interest
We have a stake on both sides of this, so both halves belong on the page. Mode sells an AI Visibility Audit and ongoing monitoring, and a finding that says one measurement is not enough is commercially convenient for anyone selling the recurring kind. It is also awkward for us, because it means our own one-off check is a sample too, and we say so in the check itself. We would rather sell you an honest instrument than a confident number. Nobody in this field can guarantee what a model will say, on any given day, and this data is a fairly direct demonstration of why.
Method, limits and how to check us
Measurements were taken through the engines’ official APIs between 13 and 30 July 2026. Cited URLs were reduced to their hostnames, with www stripped, and consecutive measurements of the same question on the same engine compared using set overlap, with containment reported alongside to control for varying citation counts. Only aggregates are published, never individual businesses. Every figure quoted above is stated in full on this page, with its sample size, so the arithmetic can be followed without access to our raw file.
The limits matter. This is overwhelmingly a Perplexity measurement, 356 of the 380 pairs. Claude contributed 24 pairs, pointing the same direction at 69% overlap, but the magnitude here is Perplexity weighted. It measures the API surface, which is not the same thing as the consumer chat apps most people actually use, so read these as trend numbers for one surface rather than as what your customer sees in the app. Seventeen days is a short window. Gemini produced no usable data in this period because of an account quota issue at our end, not a fault of the engine. And most of the question bank runs three times a week, so a consecutive measurement is one to two days apart rather than exactly 24 hours.
We will re-run this at around 90 days. If the overlap climbs materially, that would suggest what we measured is an early market settling down, and we will update this page to say so.