September 2026: I Compared Prompt Tracking Across 8 AI Visibility Tools
Blog SEO Extension Revenue
PRESS
BBC Forbes The Guardian TechCrunch Bloomberg
Written by Glen Allsopp | +322 this month

I Tracked 70K+ ChatGPT and AI Overview Responses to Measure Their Volatility

A look at 2,600 total responses checked daily, over 28 days.

I Tracked 70K+ ChatGPT and AI Overview Responses to Measure Their Volatility

It's no secret that asking the same question to the same AI platform, seconds apart, can produce both different brand recommendations and cited links.

But just how volatile are the results, and what, if anything, remains consistent?

I first researched this topic in February, when I analyzed the same prompts hundreds of times. Exact matches were rare, but the most prominently named brands appeared frequently. 

I wanted to see if that held up today, checking over a longer timeline. 

The Dataset

In this new research, I've looked at over 70,000 responses, checking 1,300+ prompts daily in both ChatGPT and AI Overviews, to see how they shift over four weeks.

I find a lot of value in larger studies, including those by my colleagues at Ahrefs, but since individual brands aren't likely to track millions of prompts on their own, I think this angle has value.

Prompts were primarily commercial, based on keywords with traditional search volume, then converted into a more conversational style. They include questions like: "What are the best binoculars for stargazing?" Best Products, What's the best digital marketing agency in Sydney?" Agency, and "I'm looking for the best 3D design software" Software.

I'm not claiming these prompts are perfect. They're not.

While I've covered in detail how to come up with prompts to track, I don't have exact data here, and real-world requests often come as follow-ups in conversations, rather than always starting from a new chat.

The prompts you track should be specific to your business and what you're focused on. 

While I don't want to downplay the effort put into this research, view this as someone actively working in this space who was curious to pull some numbers for themself.

I'm always trying to learn and expand my knowledge, and sharing data like this is one way I do that. 

The top brands for a prompt show up on most days

For a typical prompt, the brand each platform names most often shows up on most days, rather than rotating in and out at random.

On ChatGPT, the leading brand appeared on 82% of days.

On Google AI Overviews, that figure is higher, at 89%.

To check this wasn't just a result of picking the winner after the fact, I also identified each prompt's leader using only the first 14 days of data. That brand still appeared on 79% (ChatGPT) and 85% (AI Overviews) of the following 14 days.

To show that visually, below, each row is a brand named on at least two days (sorted by most-present first). Each column is a day, with the oldest on the left:

ChatGPT

Sage Intacct 27

Brands named on two or more days, aligned with the other engine.

AI Overviews

Sage Intacct 21

Brands named on two or more days.

  • Named in the top 3
  • Top 10
  • Lower
  • Not named
  • No answer collected

One row per brand, one column per day, oldest on the left. Checked over a 28-day period.

The top spot is close. The gap between the first and second brand is typically around two days out of 28, and the leading brand changed between the first and second fortnight on roughly half of prompts.

For example, for a prompt asking for advice on the best financial services software on ChatGPT, Sage Intacct appeared on 27 of the 28 days. No other brand appeared on more than 18 days.

18% of ChatGPT prompts had at least one brand named on every single day, compared to 29% of AI Overviews. 

The "core" of a response barely moves, but the tail turns over daily

For each prompt I sorted every brand named over the 28 days into two groups: core brands, named on at least 80% of days, and the tail (everything else). Then I compared each day's list with the previous day, counting both brands that dropped out and those that newly appeared.

On ChatGPT, the core brands changed by 13% from one day to the next, against 78% for the tail. On AI Overviews, they changed by 11% and 65%.

Cited domains followed the same pattern: 14% vs 74% on ChatGPT, and 12% vs 61% on AI Overviews.

Domains

Core domains

14% changed

Tail domains

74% changed

Brands

Core brands

13% changed

Tail brands

78% changed

  • Stayed the same
  • Changed

Because core is defined by how often a brand appears, I also tried sorting brands using only the first two weeks of data, then scoring the second two. The core changed a bit more (20% on ChatGPT, against 71% for the tail), but the gap was still clear.

The catch is that the core is small. Only about half of the ChatGPT prompts that named brands had a core brand (around 6 in 10 on AI Overviews), and when a core existed, it was usually just one or two names.

If you recheck a prompt response, most of the names will change, but the few named most often rarely drop out.

Cited URLs change frequently, but their type rarely does

Both ChatGPT and AI Overviews lean heavily on articles, which account for around 6 in 10 citations. For the rest, they differ.

ChatGPT cites category and listing pages, homepages, and documentation. AI Overviews cite YouTube and forums, with video making up 11% of its citations, compared to almost none on ChatGPT.

The individual pages cited change constantly. Only about a quarter of the pages ChatGPT cited for a prompt were cited again the next day. On AI Overviews, it was around half.

  • Article
  • Listing (e.g. categories)
  • Site page (e.g. About)
  • Landing page
  • Document
  • Video
  • User content
  • Tool
  • Other

The outlier on the AI Overviews tab was due to a spike in homepages appearing. The run was complete and error-free, so this looks like a one-day change in what AI Overviews cited, rather than a collection problem, though I can't rule that out. I've left it in for transparency.

On a typical day, only about 2 points of ChatGPT's mix shift, and about 1 point on AI Overviews.

The one real trend was the presence of video links in AI Overviews, which declined from about 14% of citations in the first week to about 8% in the last.

The usual leader is named first about half the time, and is in the answer more often than that

For each prompt, I took the brand most often named first and looked at how often it held that spot. 

On ChatGPT, it was named first on 50% of days and appeared somewhere in the answer on 67%. On the days it wasn't first, it usually hadn't slipped down the list, but was actually missing entirely.

On AI Overviews, the usual leader was named first on 60% of days and appeared in the answer on 78%. When it wasn't first, it was missing just over half the time.

This isn't always the same brand as the leader in the earlier section, which was the brand that appeared on the most days. They match on only about half of prompts, which is partly why the figures are lower.

The graphic below shows this for each category:

  1. Agency 41%
  2. Beauty 50%
  3. Best Products 41%
  4. Fashion 40%
  5. Finance 45%
  6. Health Not applicable
  7. Marketing 63%
  8. Software 64%
  9. Tech Stack 67%
  • Named first
  • In the answer, not first
  • Not mentioned

For one prompt, ChatGPT named a different agency first on every single day of the month.

At the other extreme, SEO Sherpa was named first on all 28 days for "What's the best SEO agency in Dubai?" in ChatGPT.

Pick two random days and you'll get an identical list of brands or cited domains less than 3% of the time

For the typical prompt, pick any two days and ChatGPT returned an identical set of brands on 0.3% of pairs, and an identical set of cited domains on 2.4%.

Google AI Overviews were similar: 1.1% for brands and 0.3% for domains.

Exact brand list match

ChatGPT 0.3%
AI Overviews 1.1%

Two random days, identical list

Exact domain set match

ChatGPT 2.4%
AI Overviews 0.3%

Two random days, identical set

This is a deliberately strict test.

If ChatGPT names 10 brands and 9 match the next day, that counts the same as if none matched.

I've run these tests before and the number still feels very low to me, so here's a better visual representation of how small changes might be, based on actual responses. 

Cream = not named the next day

Day 1

  • Refine Labs
  • Powered by Search
  • Directive Consulting
  • Demandbase
  • New Breed
  • Kalungi

Day 2

  • Powered by Search
  • Directive Consulting
  • Refine Labs
  • Obility
  • Demandbase
  • Ironpaper
  • Velocity Partners
  • Kalungi

Day 3

  • Powered by Search
  • Directive Consulting
  • Refine Labs
  • Kalungi
  • New Breed
  • SmartBug Media
  • Demandbase
  • Ironpaper

41% of ChatGPT prompts never returned the same set of brands twice, not even on one pair of days in the month. On AI Overviews, that number was 30%.

Averages are higher, where exact citations show an average of 3.2% (ChatGPT) and 1.9% (AI Overviews) of the time, as a small number of promptsβ€”particularly in Health and Marketingβ€”repeat citations far more often.

The typical prompt is represented above. 

The most prominent domain is cited most days, but is cited first less than half the time

For each prompt, I took the domain each platform cited most often. For a typical prompt, it was cited in 82% of daily answers on ChatGPT and 89% of daily AI Overviews.

Being cited first is a different story. That same domain was the first citation on only 43% of days on ChatGPT, and 28% on AI Overviews.

Each column below is one prompt, and each row is one day (oldest at the top):

Leading domain cited 82%
Leading domain cited first 43%

AI Overviews are more likely to cite the leading domain somewhere in the answer, but less likely to put it first.

The consistency of AI answers, unsurprisingly, depends on the topic you're asking about

The table below brings these measures together by category. Each prompt is scored individually, across its daily answers.

The bold figure is the typical prompt (the median), and the brackets show the average.

The topic you're researching makes a big difference.

Tech Stack and Marketing prompts named the same first brand around half the time, while in the Fashion category that happened only 12% of the time.

Health questions very rarely named brands, so I left those sections blank, but Health had the most consistent citations of any category, repeating the same first domain on more than half of days.

Category First brandBrands + orderBrands (any order)LinksFirst domainDomains
Agency 20.5% (avg. 25.9%) 0% (avg. 0.1%) 0% (avg. 0.4%) 0.3% (avg. 1.2%) 30.3% (avg. 33.9%) 1.9% (avg. 4.4%)
AI Visibility 31.1% (avg. 34.8%) 0% (avg. 0.6%) 0.5% (avg. 1.2%) 0% (avg. 0.1%) 17.4% (avg. 27.5%) 0% (avg. 6.1%)
Beauty 29.1% (avg. 35.8%) 0% (avg. 0.6%) 0.3% (avg. 1.8%) 0.5% (avg. 1.2%) 31.0% (avg. 36.4%) 2.6% (avg. 4.8%)
Best Products 18.9% (avg. 24.8%) 0% (avg. 0.3%) 0% (avg. 0.6%) 0.5% (avg. 1.7%) 29.4% (avg. 34.8%) 2.5% (avg. 5.1%)
Fashion 12.3% (avg. 22.7%) 0% (avg. <0.1%) 0% (avg. 0.1%) 1.1% (avg. 2.1%) 23.1% (avg. 27.5%) 3.0% (avg. 4.2%)
Finance 24.0% (avg. 30.8%) 0.3% (avg. 0.9%) 0.5% (avg. 2.2%) 0% (avg. 0.5%) 23.7% (avg. 29.7%) 1.3% (avg. 5.2%)
Health β€” β€” β€” 2.0% (avg. 3.5%) 55.8% (avg. 59.2%) 15.3% (avg. 19.1%)
Marketing 49.5% (avg. 58.0%) 1.1% (avg. 2.2%) 3.7% (avg. 5.6%) 0.3% (avg. 0.6%) 46.8% (avg. 51.8%) 12.6% (avg. 16.7%)
Software 44.6% (avg. 51.5%) 0.3% (avg. 1.0%) 1.6% (avg. 3.8%) 0% (avg. 0.2%) 24.1% (avg. 26.1%) 0.7% (avg. 2.0%)
Tech Stack 56.7% (avg. 61.6%) 1.1% (avg. 2.2%) 3.4% (avg. 5.9%) 0% (avg. 0.5%) 26.5% (avg. 29.1%) 2.1% (avg. 3.5%)
All categories 29.9% (avg. 36.6%) 0% (avg. 0.7%) 0.3% (avg. 2.1%) 0.3% (avg. 1.1%) 29.9% (avg. 36.3%) 2.4% (avg. 7.2%)

Only prompts that named at least one brand on a quarter of days or more are included, so prompts that never name anything don't look perfectly consistent.

There's an important caveat: Some brand volatility is simply how brands are referenced

AI answers often name a company in different ways.

Sometimes the same prompt will answer naming "Ahrefs" and other times "Ahrefs Brand Radar".

HubSpot might also be referred to as "HubSpot CRM", and in more subtle cases, my friend Benji's agency can be called "Grow and Convert" or "Grow & Convert".

You might see the same vacuum brand recommended, but a different product of theirs can show up. 

Throughout this study, I counted names exactly as AI wrote them. That's deliberate, as it's what the answers actually say, and automatically matching variations will always get some wrong. 

Not all recommendations are created equally, which can also influence results.

For example, you can have an entire bullet point dedicated to recommending you, or get thrown in at the end of a response, referenced in a way like, "Try extensions like the Ahrefs Toolbar/Detailed SEO Extension if you want single page checks while browsing".

I tested how much these variations matter, starting with manually combining the name and product variants of 55 brands (411 names in total).

For these brands specifically (not all), 36% of apparent drop-offs on ChatGPT and 30% on AI Overviews were the same company under a different name the next day. 

I then tried automating the process across the entire dataset: The usual leader was named first 5–8 points more often, but the tail barely changed.

Most of it is genuinely different brands, and identical lists still rarely repeated.

As you can imagine, this also varies a lot by the company being recommended.

Most of HubSpot and Wells Fargo's initial drop-offs were due to renaming, yet brands like Anaplan and Freshdesk were almost always named the same way. 

If you're tracking your own brand, make sure you're accounting for name and product variations. Your inclusion in a response may be more stable than your tracking suggests.

Should you be tracking multiple platforms?

So far we've covered whether the same recommendations appear one day to the next.

A different question is whether ChatGPT and AI Overviews recommend the same things in the first place.

Mostly, they don't.

For the same prompt, the two platforms agreed on the most-present brand only 27% of the time.

They also lean on different sources, and in different amounts.

Even the leader on one platform often isn't a regular on the other.

In only about half of prompts did ChatGPT's leading brand appear on most days in AI Overviews, and the same was true the other way round.

Despite the differences, I don't want to say every brand should check their visibility in every platform, just as I don't recommend people track every possible prompt and variation.

Personally, I'm most interested in traditional Google search rankings, ChatGPT responses, AI Overviews, and Claude responses.

I'll probably place more importance on AI Mode going forward, and I'm also testing a few lesser-known sources I'm really fascinated by, which I'll share soon. 

Even the most prominent domains can differ a lot

Before the following lists confuse you, as they don't all typically feature on "most-cited" domain studies you might have seen, remember this is just for the specific prompts and categories I was tracking.

Here are the top 20 most cited domains (including subdomains where relevant) across each platform:

Cream = cited on both platforms

ChatGPT

# Domain DR Cited in
1 clutch.co 91 3,424
2 ahrefs.com 91 2,517
3 developers.google.com 99 2,332
4 nerdwallet.com 90 1,874
5 mayoclinic.org 92 1,770
6 semrush.com 92 1,646
7 capterra.com 91 1,640
8 forbes.com 94 1,625
9 allure.com 87 1,414
10 softwareadvice.com 87 1,251
11 wsj.com 92 1,224
12 techradar.com 91 1,167
13 learn.g2.com 91 1,108
14 bankrate.com 90 1,061
15 cdc.gov 93 1,014
16 designrush.com 90 1,014
17 g2.com 91 1,009
18 vogue.com 91 982
19 medlineplus.gov 91 970
20 support.google.com 99 866

AI Overviews

# Domain DR Cited in
1 youtube.com 99 12,299
2 reddit.com 95 8,216
3 semrush.com 92 1,809
4 mayoclinic.org 92 1,799
5 my.clevelandclinic.org 92 1,698
6 nerdwallet.com 90 1,612
7 zapier.com 91 1,534
8 clutch.co 91 1,459
9 forbes.com 94 1,182
10 quora.com 91 1,090
11 bankrate.com 90 1,053
12 allure.com 87 910
13 developers.google.com 99 786
14 cnbc.com 92 778
15 linkedin.com 99 759
16 ahrefs.com 91 730
17 facebook.com 100 700
18 byrdie.com 85 693
19 designrush.com 90 669
20 g2.com 91 651

In my research, ChatGPT and Google's AI Overviews share 18 of the top 50 domains (including subdomains), 38 of the top 100, and 169 of the top 500.

ChatGPT named more brands. AI Overviews cited more sources per answer.

On answers that named at least one brand, ChatGPT named an average of 5.9 to AI Overviews' 5.1.

Brands per response

ChatGPT 5.9 +0.1% vs prior week
AI Overviews 5.1 +0.4% vs prior week

Answers that named at least one brand

Citations per response

ChatGPT 5.5 +6.4% vs prior week
AI Overviews 6.8 +7.7% vs prior week

Answers that cited at least one link

Domains per response

ChatGPT 3.7 +2.4% vs prior week
AI Overviews 6.0 +7.3% vs prior week

Registrable domains behind those links

Citations went the other way: AI Overviews averaged 6.8 links from 6.0 domains per answer, against ChatGPT's 5.5 links from 3.7 domains.

Final thoughts

As I've said before, your best insights into volatility are going to come from tracking your own space. 

How AI cited domains in Health, for example, was much more stable than in an industry like Fashion.

This shouldn't be a surprise to anyone, but don't judge results from a single day, even if your brand isn't mentioned (or site isn't cited).

The aim, if you have one, is to be part of the smaller, stable core, focusing more on overall presence compared to how often you're the first recommendation.

Make sure you're counting every way your brand gets named, including different products you might offer. 

I like to monitor multiple platforms as they clearly reward different brands and sites, but you don't have to go overboard here. 

When I personally monitor AI responses, I want to know things like:

  • Whether newly-published content starts making its way into an answer
  • Topics we've covered well but aren't being cited for
  • The most commonly cited domains and URLs, especially new(er) ones which later stabilize
  • What AI says about us directly, including without relying on web search
  • Inaccuracies around pricing, features (including new ones), and who we serve that has a relevant citation
  • What happens during any clear shifts, such as model updates

I could go on, but I think you get the picture. 

How far you take prompt tracking and how much "weight" you give to responses depends on among other things, the niche you're in and the ROI you can potentially attribute to AI and similar channels. 

I find a lot of value in it, but I also understand the volatility in responses, prioritize taking action on what I see, and don't push monitoring to the extreme.

If you have an Ahrefs subscription, a number of prompts are included to help you get started. Though I mentioned this in the intro, I should disclose again that I work at Ahrefs.

You don't need an Ahrefs subscription to track custom prompts, which actually makes Ahrefs the most competitive offer in some recent research I conducted. 

As I've said a few times, if you're going to start monitoring AI visibility, start small and focus on phrases where you're likely to take action based on your findings.

You can always expand or switch your focus later on. 

Methodology, and opportunities for improvement

Each prompt was checked once per day from the US using a non-logged-in account from the 26th of August to 22nd of September.

While they're increasing in prominence across search results, AI Overviews don't always appear for a term. 

Brands in responses were extracted using a custom, fine-tuned OpenAI model, and samples were checked manually for accuracy.

Twenty-eight days is long enough to spot trends and core brands in your space, but there are always different things I can experiment with on the testing side.

In future iterations of this, I would like to try:

  • More platforms. Claude and Google's AI Mode are interesting to me, though I know others would like to include Perplexity in there.
  • Commercial prompts only. The set mixes some informational queries with buying ones. I'm curious whether any sites become leaders on the informational side, and that's how I think many brands will view their visibility, but I should also test by focusing solely on commercial questions. 
  • A longer window. This one should be pretty self-explanatory, but would also be great to help point out specific model shifts. 
  • Combine even more brand and product variations. Automation was not as effective as a manual check, and how a brand is represented also matters. In a future test, I would map these to more individual brands to see how core stability improves further. 

I've made every effort to ensure all calculations were accurate and made sense, with multiple reviews of the data.

That said, this was one of the more challenging studies I've taken on (if not the most, surprisingly), so something may have been overlooked. Take the research in this view accordingly, but know it will be updated if any oddities are found.

If you would like to discuss this article, I would love your thoughts over on LinkedIn or X (those are direct post links). 

Written by Glen Allsopp, the founder of Detailed, which is now (very proudly) an Ahrefs company. You may know me as 'ViperChill' if you've been in internet marketing for a while. We're behind the Detailed SEO Extension for Chrome & Firefox, which currently has 600,000 weekly users. Clicking the + button tells us what you enjoy reading. Social sharing is appreciated (and always noticed). You can also follow me on Twitter and LinkedIn.

"Think of us like Bloomberg for Digital Goliaths."

Exclusive insights from tracking the traffic & revenue of 3,078 digital goliaths.

β€œGlen found a very sneaky technical SEO issue on our homepage. Sometimes a fresh set of eyes goes a long way.”

Bill King

β€œGlen's recommendations helped us improve crawl budget, remove deadweight pages and led to overall improvements in organic traffic to our key pages.”

Steve Toth

β€œI've been a practitioner of digital marketing for over a decade and I've learned more from Glen about SEO than anyone else.”

Clay Collins