AEO, GEO and SEO research in 2026. What AI actually cites.
The AI Visibility Report: what ChatGPT, Perplexity, Gemini, Google AI Mode and AI Overviews actually cite, measured on 7,949 prompts across three waves.
7,949
prompts, 158 categories
5
AI engines measured
3
collection waves, July to September
39,572
eligible answers in the final wave
239,841
citations, 27,765 domains
1.18M
citation rows in the archive
Research edition: July, August and September 2026 · OMNIFAMOUS Research
AEO, GEO and SEO Research 2026
The AI Visibility Report
Jorge Ferreiro, OMNIFAMOUS. Published 2026-09-06. Every figure names its cohort and reproduces from one frozen snapshot.
7,949
prompts, 158 categories
5
AI engines measured
3
collection waves, July to September
39,572
eligible answers in the final wave
239,841
citations, 27,765 domains
1.18M
citation rows in the archive
Executive summary
A prospect asks an AI assistant which product to buy. Your company appears in the answer, but the link goes to a comparison site. A second prospect asks a similar question and sees a competitor. Your team checks again, gets a third answer, and starts arguing about whether the content strategy is working. This report exists so that argument can be settled with evidence instead of screenshots.
Between 29 July and 5 September 2026 we submitted the same 7,949 prompts, drawn from 158 categories and 29 question archetypes, to five AI answer surfaces: ChatGPT, Perplexity, Gemini, Google AI Mode and Google AI Overview. We did it three times, roughly three weeks apart, and beside two of the waves we ran a 1,000-prompt repeat panel so that every cross-wave change could be compared with how much each engine disagrees with itself on the same day. The final wave closed with 39,572 eligible answers, 239,841 citations and 27,765 distinct cited domains. The archive behind the study holds 125,863 answer rows and 1,181,828 citation rows.
Take the full report with you, or keep reading below.
Key findings
Ten numbers, each with its caveat attached
Every figure below clears the engine's own same-day noise floor or says why it cannot. Nothing here is a causal claim.
01
14.8% to 64.6%
of July's cited domains survived to September
AI Mode kept 14.8%, ChatGPT 20.4%, Gemini 29.5%, AI Overview 42.3%, Perplexity 64.6%. Every figure clears that engine's own same-day noise floor.
02
5.40 from 15.51
sourced pages per ChatGPT answer, August to September
AI Mode fell from 17.31 to 3.34 on the same prompts. Perplexity and Gemini barely moved. A vendor interface change sits inside the ChatGPT interval.
03
44.9% to 2.7%
of ChatGPT answers citing Reddit, July to September
AI Mode fell from 60.2% to 21.5%. Gemini did not move. AI Overview rose, then fell below July. A two-point trend predicted the third point wrong twice.
04
80.0% to 0.0%
of ChatGPT answers carrying an ad, August to September
Confirmed by two independent tests. The surface that is growing is Google AI Mode: 0.21% of answers in July, 8.94% in September.
05
64.1%
of missing AI Overviews appeared on a same-day retry
237 of 370 empty captures returned an AI Overview minutes later. A once-a-day tracker reports two thirds of those absences as fact.
06
42.07%
of AI citations point to brand-owned domains
Review publishers take 17.60%, social video 7.32%, forums 5.31%, news 2.03%. 57.72% of cited pages were not on Google's page one for the same query.
07
5.61% to 34.66%
of Google page-one results an engine also cites
ChatGPT cites 5.61% of page one, AI Overview 34.66%. On ChatGPT, 70.81% of citations go to a domain that was not on page one at all.
08
55.6% vs 25.5%
page-one written pages vs videos cited by any engine
Written pages yield 1.9 answers per asset, YouTube videos 1.37. ChatGPT and Perplexity nearly stopped citing YouTube; Google surfaces did not.
09
9 of 10
local-intent answers relocated by a city target on AI Mode
Against 1 and 2 of 10 in the control arms. On brand and product prompts, geography changed nothing: 0 of 30 in every arm on every engine.
10
9.76% or 38.25%
brand flip rate on ChatGPT, depending on the denominator
The same 8,429 transitions divided by 86,393 opportunities or by the 22,037 where the brand ever appeared. Neither is a customer's risk of losing a mention.
A prospect asks an AI assistant which product to buy. Your company appears in the answer, but the link goes to a comparison site. A second prospect asks a similar question and sees a competitor. Your team checks again, gets a third answer, and starts arguing about whether the content strategy is working. This report exists so that argument can be settled with evidence instead of screenshots.
Between 29 July and 5 September 2026 we submitted the same 7,949 prompts, drawn from 158 categories and 29 question archetypes, to five AI answer surfaces: ChatGPT, Perplexity, Gemini, Google AI Mode and Google AI Overview. We did it three times, roughly three weeks apart, and beside two of the waves we ran a 1,000-prompt repeat panel so that every cross-wave change could be compared with how much each engine disagrees with itself on the same day. The final wave closed with 39,572 eligible answers, 239,841 citations and 27,765 distinct cited domains. The archive behind the study holds 125,863 answer rows and 1,181,828 citation rows.
The headline result is turnover. On every one of the five surfaces, the set of domains an engine cited in July had substantially changed by September, and on every surface that change is larger than the engine's own same-day variability. AI Mode retained 14.8% of its July domains, ChatGPT 20.4%, Gemini 29.5%, AI Overview 42.3% and Perplexity 64.6%. Two surfaces also cut the number of sources they expose per answer by roughly two thirds. ChatGPT's ad unit, present in 80% of August answers, was absent from every September answer. And two thirds of the AI Overviews that were missing on first capture appeared when the same query was resubmitted minutes later.
The second result is about ownership. Brand-owned domains are the single largest class of cited source (42.07% of citations), review publishers second (17.60%), and 57.72% of cited pages were not on Google's first page for the same query. What AI cites and what Google ranks are related but different inventories, and the difference is largest on ChatGPT, where 70.81% of citations go to a domain that was not on page one at all.
The third result is methodological, and it is the one we would ask you to take away if you take away only one thing. A single capture of an AI answer is an observation, not a ranking. Retention, reach, share and availability each need a stated denominator and a stated retry policy, and the same brand transitions can be reported as 9.76% or 38.25% depending on which one you pick. Chapters 11 to 14 turn that into a measurement worksheet, a diagnostic tree, a content repair brief and a 30-day operating plan you can run on your own categories.
1. What we measured and how
An engine name is shorthand for a measurement surface. Every capture in this study recorded the provider, interface, country (US), language (English), request settings and, where the provider supplied it, the serving model. Distinct provider task IDs establish separate requests; they do not establish independent generation or the absence of provider caching, and this report never assumes either.
Eligibility requires a completed task and a non-empty answer body. Copilot is excluded from every cross-engine comparison because its prompt coverage is incomplete (954 July requests only). Retries of the same provider task are not new observations. A failed task is missing evidence, not a zero-brand answer.
Collection outcomes by run
Run
Requested keys
Eligible answers
AI Overview eligible
Every other surface
study-2026-07-29
39,745
39,546
7,765 / 7,949 (97.68%)
7,937 to 7,949 of 7,949
study-2026-08-19
33,000
32,492
6,112 / 6,600 (92.61%)
6,580 to 6,600 of 6,600
panel-2026-08-19
5,000
4,933
933 / 1,000 (93.30%)
1,000 of 1,000
study-2026-09-05
39,745
39,572
7,776 / 7,949 (97.82%)
7,949 of 7,949
panel-2026-09-05
5,000
4,968
968 / 1,000 (96.80%)
1,000 of 1,000
Requested keys are distinct run-prompt-engine identities. The September AI Overview figure includes 196 answers obtained by a same-day resubmission; excluding them, September AI Overview eligibility is 7,580 of 7,949 (95.36%).
The analytic samples behind each comparison
Analysis
Question it answers
Size
Full September description
What did the engines return across the whole prompt set?
39,572 eligible answers
Three-wave complete-case comparison
What changed on the same prompt-engine keys in all three waves?
32,326 keys
Expanded July-to-September comparison
What changed over the longer interval, including prompts outside the August cohort?
39,405 eligible pairs
Repeated September panel
How much do two same-day captures of the same prompt differ?
4,958 eligible pairs
These cohorts have different sizes and must not be interchanged in a headline. "We collected nearly forty thousand requests" does not mean a given comparison contains forty thousand matched pairs.
The prompt set was built from customer language rather than from keyword tools: sales notes, support questions, on-site search and interviews, with synthetic variations labelled separately. Each prompt is tagged with the decision it represents (category discovery, fit, comparison, pricing, objection, integration, switching, implementation). A prompt portfolio has a purpose; it is not a pile of phrases a company happens to perform well on.
2. The cited web turned over in four months
This is the result the study was collected to produce. On every surface, the set of domains an engine cited in July had substantially changed by September, and on every surface the change is larger than that engine's own same-day disagreement with itself.
Cited-domain retention, July to later waves, with each engine's same-day floor
Surface
July to August (pairs)
July to September (pairs)
Same-day floor
Clears the floor
ChatGPT
27.7% (6,529)
20.4% (7,865)
52.0%
yes
Perplexity
71.0% (6,511)
64.6% (7,847)
94.0%
yes
Gemini
29.3% (6,189)
29.5% (7,495)
43.3%
yes
Google AI Mode
32.5% (6,567)
14.8% (7,935)
41.3%
yes
Google AI Overview
49.4% (5,918)
42.3% (7,539)
63.1%
yes
Cohorts: July to August is the matched main-wave cohort (32,361 eligible pairs over 6,600 prompts). July to September is the expanded cohort (39,405 eligible pairs over all 7,949 prompts). Retention is the share of a July capture's cited domains that reappear in the later capture, averaged over answer pairs; pairs with an empty July set are excluded, which is why pair counts differ.
Like-for-like series on the three-wave complete-case cohort
Surface
Pairs
July to August
July to September
August to September
ChatGPT
6,590
27.7%
20.8%
37.0%
Perplexity
6,600
71.0%
64.6%
74.5%
Gemini
6,598
29.3%
29.7%
33.6%
Google AI Mode
6,579
32.5%
15.0%
15.4%
Google AI Overview
5,959
49.4%
42.9%
44.7%
Two engine-specific readings survive. Gemini's July-to-September retention (29.7%) is indistinguishable from its July-to-August retention (29.3%): whatever turnover happened, happened by August and did not compound. AI Mode is the opposite: its August-to-September retention (15.4%) is barely above its July-to-September retention (15.0%), so almost all of AI Mode's turnover occurred in the second, shorter interval.
How the noise-floor gate works
Retention and Jaccard are similarity measures. A cross-wave value below the floor means the two waves agree with each other less than the engine agrees with itself on the same day, which is exactly what makes the change real rather than noise. The floor used is the August panel, which is the stricter of the two available: September's panel agreed with itself more on every surface, so gating against August makes the finding harder to obtain, not easier.
Same-day self-agreement, broad citation policy (denominators in brackets)
Surface
August panel retention
September panel retention
August Jaccard
September Jaccard
Perplexity
94.0% (1,000)
98.0% (1,000)
0.911 (1,000)
0.975 (1,000)
Google AI Overview
63.1% (898)
70.0% (919)
0.474 (902)
0.552 (927)
ChatGPT
52.0% (998)
53.9% (996)
0.331 (1,000)
0.398 (998)
Gemini
43.3% (774)
50.0% (896)
0.280 (884)
0.349 (953)
Google AI Mode
41.3% (961)
53.0% (882)
0.264 (996)
0.370 (948)
The September panel was drained together with the main wave (median fetch gap 0.000 hours, p90 at most 0.012 hours). That removes the August panel's ordering problem and does nothing about caching: two tasks retrieved within the same second are exactly the case where a provider cache would be least visible.
3. Two engines now expose far fewer sources
Retention fell while the number of exposed sources fell with it. On the three-wave complete-case cohort, the mean count of distinct sourced pages per answer moved as follows.
Distinct sourced pages per answer, three-wave cohort
Surface
Pairs
July
August
September
ChatGPT
6,590
12.70
15.51
5.40
Perplexity
6,600
9.87
10.03
10.06
Gemini
6,598
3.27
2.47
2.79
Google AI Mode
6,579
17.49
17.31
3.34
Google AI Overview
5,959
9.64
10.07
7.31
Perplexity and Gemini barely moved. ChatGPT and AI Mode, the two surfaces with the largest retention falls, are also the two that ended September exposing roughly a fifth to a third as many sourced pages per answer as in August. "The engines changed which sources they cite" and "the engines expose fewer sources" are both consistent with these numbers, and this corpus cannot separate them. A retention figure computed against a shrinking target set will fall even if the surviving choices are unchanged in kind.
For ChatGPT there is a documented capture-side explanation. The collection vendor's changelog records that between 25 and 26 August its mobile-web format began serving 98% of ChatGPT responses, with metadata fields potentially empty, and that from 28 August a legacy flag requests the older desktop interface on a best-effort basis. Our own side-by-side test on 5 September (20 prompts per arm) found the current default recovered a mean of 0.1 fan-out search queries per answer against 1.8 on the legacy desktop arm, with a different response key set and a different mean answer length (3,690 against 2,688 characters).
4. Reddit and YouTube: a two-point trend was wrong twice
This is the strongest available argument against building a content programme on one wave's source leaderboard. Reach counts an eligible answer once when the domain appears in its sourced evidence. All three columns are the same prompt-engine keys.
Reddit reach by wave, three-wave cohort
Surface
Pairs
July
August
September
July to September change [95% CI]
ChatGPT
6,590
44.90%
8.01%
2.69%
-42.22 pp [-43.43, -41.00]
Perplexity
6,600
27.58%
23.65%
22.71%
-4.86 pp [-5.71, -4.02]
Gemini
6,598
8.99%
8.84%
9.05%
+0.06 pp [-0.75, +0.87]
Google AI Mode
6,579
60.24%
42.48%
21.46%
-38.77 pp [-40.13, -37.42]
Google AI Overview
5,959
49.94%
52.53%
39.47%
-10.47 pp [-11.91, -9.04]
YouTube reach by wave, three-wave cohort
Surface
Pairs
July
August
September
July to September change [95% CI]
ChatGPT
6,590
3.02%
4.66%
0.20%
-2.82 pp [-3.24, -2.40]
Perplexity
6,600
28.88%
1.65%
0.59%
-28.29 pp [-29.38, -27.20]
Gemini
6,598
3.06%
2.18%
2.23%
-0.83 pp [-1.34, -0.33]
Google AI Mode
6,579
56.70%
58.46%
40.14%
-16.55 pp [-18.00, -15.10]
Google AI Overview
5,959
58.73%
60.18%
55.48%
-3.26 pp [-4.61, -1.91]
Intervals use a prompt-cluster, ratio-of-sums normal approximation. They do not adjust for category-level shocks, multiple comparisons, measurement error or changes to the provider surface.
Four readings, in decreasing order of confidence. Perplexity's YouTube reach fell by an order of magnitude in the first interval and stayed down. Gemini did essentially nothing on either domain across four months, so "AI cites Reddit less now" is not a statement about AI. AI Overview's Reddit reach rose from July to August and then fell below its July level, so a team that had extrapolated the first interval would have been wrong about the second. And ChatGPT's falls, on both domains, are the ones to trust least, because they sit inside the documented interface change and coincide with its sourced-page count dropping from 15.51 to 5.40.
What none of this supports is an instruction. It does not show that forum participation stopped working, that video stopped working, or that either will still be down next month. It shows that the composition of the cited web on these surfaces moved a long way in four months, which is an argument for measuring your own categories repeatedly rather than for reallocating a budget on the strength of a single leaderboard.
5. ChatGPT ads went to zero. AI Mode ads are growing
Advertisements create a measurement problem because they may arrive as separate structured data, appear in rendered prose, or both. A paid slot beside an answer is not an organic recommendation, and a rendered ad can inject a competitor name into organic mention detection if the two are not separated at parse time.
Share of eligible answers carrying at least one detected ad, three-wave cohort
Surface
Pairs
July
August
September
August to September change [95% CI]
ChatGPT
6,590
69.74%
80.02%
0.00%
-80.02 pp [-80.98, -79.05]
Google AI Mode
6,579
0.21%
4.38%
8.94%
+4.56 pp [+3.79, +5.33]
Google AI Overview
5,959
0.25%
0.49%
0.35%
-0.13 pp [-0.36, +0.09]
Perplexity
6,600
0.00%
0.00%
0.00%
+0.00 pp
Gemini
6,598
0.00%
0.00%
0.00%
+0.00 pp
The ChatGPT zero is not a parsing gap and not a rounding artefact. The structured ad table holds 10,380 ChatGPT rows for August and zero rows of any provenance for September, while AI Mode and AI Overview ad rows kept flowing in the same September run. Sampling 200 raw payloads per wave, 171 of 200 August payloads (85.5%) carried a non-empty ads array against 0 of 200 in September.
Two independent tests rule out the two obvious measurement explanations. The vendor's older desktop interface, requested on the same 20 prompts on the same day, returned zero ads on 18 of 18 completed tasks. And chatgpt.com itself, checked directly in a logged-out browser on 5 September, served no ad markup on three commercial prompts, two of which had carried 15 ad slots each in August.
Meanwhile the surface that is quietly growing is Google AI Mode, from 0.21% of answers in July to 8.94% in September on the same 6,579 pairs. Its ad provenance is structured in all three waves and all three were parsed under the same harmonised parser, so the specific confound that destroys the ChatGPT July-to-August comparison does not apply. On the evidence available it is the paid surface worth watching next.
One correction to our own earlier work belongs here. Regenerating the July analysis against the frozen snapshot showed that its headline claim, that ChatGPT was the only engine serving this ad unit, was already false in July: 17 AI Mode answers and 22 AI Overview answers carried a structured ad in the July run, on a rounding-to-zero scale that an array-only parser returned as literal zero.
6. An absent AI Overview is often a property of the capture
The September wave closed with 383 non-completed tasks: 370 empty, every one of them AI Overview, and 13 failed. All 383 were resubmitted the same day, unchanged.
Outcome of the same-day retry on 370 empty AI Overview captures
Outcome
n
Share of the 370 empties
Returned an AI Overview
237
64.1%
Still empty
133
35.9%
All 13 failures also recovered. The 133 residual empties reconcile as 113 on the main wave and 20 on the panel.
Two thirds of the queries that had "no AI Overview" produced one minutes later, on the same query, same country, same day. Three consequences follow, and the third is the commercial one.
A prevalence claim has to state its retry policy. "AI Overview appears on X% of prompts", measured from a single capture, understates X by roughly the share of transient empties, which here was 64% of observed empties.
A cross-wave coverage comparison is confounded unless both waves retried the same way. September's AI Overview eligibility of 97.82% against August's 92.61% cannot be read as Google expanding coverage. Excluding September's retry-obtained captures brings it to 95.36%, still above August and still below July's 97.68%, which is a mixed picture rather than a trend.
A once-a-day monitoring product inherits this directly. It captures once per prompt per day and reports an absent AI Overview as an absence. On this evidence roughly two in three of those absences would not survive a second look taken minutes later.
7. Who owns the cited web
Every cited domain in the September wave was classified by owner type using a layered scheme: manual overrides, curated exact-domain lists, TLD rules, host aliasing, a seed-brand match and evidence-based classification from page paths and titles. 9,119 of 27,765 domains were classified (32.84%), covering 86.77% of citation volume. The long tail is real: 18,646 unclassified domains hold 13.23% of citations.
Citations by owner type, September wave (39,572 eligible answers)
Owner type
Domains
Citations
Share of citations
Answers reached
Share of answers
Brand-owned
5,539
100,897
42.07%
28,950
73.16%
Review publisher
1,497
42,211
17.60%
18,800
47.51%
Unclassified
18,646
31,725
13.23%
16,407
41.46%
Social video
12
17,551
7.32%
8,937
22.58%
Other
1,710
15,635
6.52%
9,571
24.19%
Forum (UGC)
19
12,726
5.31%
8,281
20.93%
Aggregator / directory
83
5,944
2.48%
3,920
9.91%
News
77
4,874
2.03%
3,362
8.50%
Retailer
59
3,235
1.35%
1,303
3.29%
Government / education / medical
96
2,698
1.12%
1,404
3.55%
Docs / code
16
2,197
0.92%
1,104
2.79%
Affiliate
11
148
0.06%
141
0.36%
The 20 domains reaching the most answers, September wave
Rank
Domain
Owner type
Answers reached
1
youtube.com
Social video
7,821
2
reddit.com
Forum
7,549
3
g2.com
Review publisher
4,795
4
capterra.com
Review publisher
1,815
5
zapier.com
Brand-owned
1,558
6
trustpilot.com
Review publisher
1,500
7
forbes.com
News
1,282
8
google.com
Other
1,248
9
medium.com
Other
1,158
10
apple.com
Other
1,116
11
github.com
Docs / code
1,008
12
sourceforge.net
Aggregator
907
13
gartner.com
Review publisher
899
14
linkedin.com
Social video
856
15
techradar.com
Review publisher
830
16
softwareadvice.com
Review publisher
792
17
facebook.com
Social video
721
18
pcmag.com
Review publisher
701
19
slashdot.org
Aggregator
606
20
amazon.com
Retailer
598
Reach counts an answer once per domain, regardless of how many times the domain appears in it. google.com is largely Play Store listings, apple.com mixes App Store with support, and amazon.com mixes retail with AWS; the taxonomy flags these wherever it matters. A leaderboard is a description of this prompt set, not a publishing strategy.
AI citations are less concentrated than Google's first page
Concentration, size-matched draws of 66,366 citations (30 draws)
Metric
AI cited (mean, sd)
Google page one (mean, sd)
Domains needed for 50% of citations
533.2 (5.9)
328.7 (5.0)
Top-10 domain share
19.58%
25.36%
Herfindahl index
76.7 (0.8)
143.9 (1.7)
Effective number of domains
130.4
69.5
Concentration by engine, full depth, September wave
Engine
Citations
Domains
Domains for 50%
Top-10 share
Gini
Google AI Mode
30,609
6,545
118
36.03%
0.707
Google AI Overview
60,071
11,252
254
29.27%
0.726
ChatGPT
46,516
9,512
522
11.93%
0.677
Gemini
22,639
7,930
738
11.57%
0.555
Perplexity
80,006
14,242
460
16.12%
0.726
The Google-owned surfaces concentrate their citations on far fewer domains than ChatGPT, Gemini or Perplexity do. AI Mode needs 118 domains to account for half of its citations; Gemini needs 738. For a brand, that is the difference between a surface where a handful of publishers carry most of the evidence and a surface where the long tail is still in play.
Vendor sites are cited, but their help centres barely are
Of the 1,832 seed vendor domains in the study, 1,690 (92.25%) were cited at least once in the September wave. The median vendor domain reached 10 answers, the 90th percentile 46, and the maximum 1,570. Of the 55,859 citations to vendor domains, 2.07% pointed at a docs subdomain, 1.86% at help, 1.44% at support, 1.11% at blog and 0.60% at status. The marketing site is what gets cited; the documentation that would actually settle a fit or integration question is a rounding error.
8. What Google ranks and what AI cites are different inventories
For 7,938 of the 7,949 prompts, the study also captured Google's first page of organic results (66,366 rows, mean depth 8.4). That allows one question per engine: of the pages Google ranked on page one for this query, how many did the engine also cite?
Share of Google page-one results also cited, by engine and rank, September wave
Engine
Share of page one cited
Cited at rank 1
Cited at rank 9
Rank-1 to rank-9 lift
Google AI Overview
34.66%
48.15%
25.23%
1.9x
Perplexity
24.53%
39.13%
14.88%
2.6x
Google AI Mode
12.79%
27.21%
5.55%
4.9x
Gemini
11.23%
15.32%
6.40%
2.4x
ChatGPT
5.61%
11.49%
2.54%
4.5x
Where an engine's citations come from, relative to Google page one
Engine
Citations
Exact page on page one
Same domain, different page
Domain not on page one
Google AI Overview
60,007
37.45%
19.56%
42.98%
Google AI Mode
30,581
27.77%
22.25%
49.98%
Gemini
22,605
32.94%
11.15%
55.91%
Perplexity
79,909
20.38%
15.16%
64.45%
ChatGPT
46,493
8.02%
21.18%
70.81%
Rank still matters inside every engine: a page at position one is 1.9 to 4.9 times more likely to be cited than a page at position nine. But the overall overlap is small everywhere except AI Overview, and on ChatGPT seven in ten citations go to a domain that Google did not rank on page one at all. Across all engines, 57.72% of cited pages were off Google's first page. The off-page-one rate is lowest for forum citations (10.40%), because Reddit ranks, and highest for the unclassified long tail (83.24%).
9. Written pages beat video as cited assets
Share of answers citing YouTube, by engine
Engine
July answers
July citing YouTube
September answers
September citing YouTube
ChatGPT
7,865
2.8%
7,885
0.3%
Perplexity
7,847
28.9%
7,944
0.6%
Gemini
7,495
3.1%
7,175
2.3%
Google AI Mode
7,935
56.6%
7,229
46.1%
Google AI Overview
7,679
60.1%
7,477
57.1%
Among assets that Google ranked on page one, a written page was cited by at least one engine 55.6% of the time, a UGC or social page 36.0% of the time, and a YouTube video 25.5% of the time. Per asset, written pages yielded a mean of 1.9 answers (over 111,261 assets) against 1.37 for YouTube videos (over 10,818). Video is a Google-surface phenomenon: AI Mode and AI Overview kept citing it in roughly half of answers while ChatGPT and Perplexity all but stopped.
A cited-video sample omits the videos that were never cited, so none of this identifies the causal effect of making a video. It does say that if your category is answered mainly by ChatGPT or Perplexity, a video-first source strategy has very little evidence behind it in this corpus.
10. Geography moves local intent, not brand visibility
Every capture in the study used a national US setting. That was an assumption, so it was tested: 40 prompts, three engines, three arms (national, an identical national repeat as the same-day floor, and a San Francisco target), 360 tasks. The prompt set was 30 corpus prompts plus 10 local-intent prompts as a positive control.
Answers naming San Francisco, the Bay Area, Oakland, Berkeley or the Mission
Engine
Mechanism
Prompt group
National
National repeat
City target
Google AI Mode
location: city
Local intent
2/10
1/10
9/10
ChatGPT
state: CA
Local intent
0/10
0/10
2/10
Perplexity
state: CA
Local intent
0/10
1/10
0/10
All three
either
Brand and product prompts
0/30
0/30
0/30
AI Mode's city target genuinely relocates the answer. A state-level proxy is much weaker, which is what a state-level proxy should be. Perplexity appears to ignore the state parameter entirely: its cited domains were byte-identical across all three arms on all 40 prompts. And on brand and product questions, geography did nothing at all.
11. Six outcomes that should never share one score
The most consequential measurement choice is deciding what counts as a success. An answer can mention your company without recommending it. It can recommend you without citing your website. It can cite your documentation without putting your product on a shortlist. A sponsored placement can sit beside the answer and create an apparent organic win. If those outcomes become one score, the score conceals the decision your team needs to make.
The six outcomes
Outcome
Definition
What it is not
Answer availability
The surface returned an eligible answer for the prompt
Whether your brand appears
Brand mention
The organic prose names your company or a verified alias
A name inside a tracking URL, image target, source-card excerpt or ad
Recommendation
The answer suggests your product for the stated use case
A negative comparison that happens to contain your name
Own-site citation
An eligible organic source points to a domain you own
An endorsement; it may support one narrow fact
Inline destination
A navigable link to your business inside the prose
A source listed in a separate citation module
Paid or commerce exposure
An ad, shopping destination or product card
An organic recommendation
A brand mentioned in 40 eligible answers and cited in 10 has two measured outcomes. Reporting a combined "50 visibility events" double-counts some observations and removes the information needed to improve the result.
A denominator can make the same brand transitions look four times larger
For ChatGPT, the seeded category-brand roster creates 86,393 eligible brand opportunities across the paired July and August prompts. There are 5,072 observed losses and 3,357 gains, totalling 8,429 flips. Dividing by all 86,393 opportunities gives 9.76%. Dividing the same 8,429 flips by the 22,037 opportunities where the brand appeared at least once gives 38.25%. Neither number is the probability that a customer account loses an AI recommendation. The first weights every category-brand opportunity, including brands absent on both captures. The second conditions on ever appearing. Both measure detected name presence, not favourable recommendation.
Every rate needs a stated population
Metric
Numerator
Denominator
Trap
Brand mention rate
Eligible answers containing the brand
Eligible answers in scope
Report availability separately
Domain reach
Eligible answers containing the domain at least once
Eligible answers
Ten citations in one answer is one reached answer
Citation-row share
The domain's eligible citation rows
All eligible citation rows
Sensitive to how many sources each answer carries
Retention
Domains in both captures
Domains in the first capture
Undefined when the first set is empty; never insert zero
Jaccard
Intersection
Union
Both-empty is its own state, not perfect stability
Recommendation rate
Labelled favourable recommendations
Eligible answers
Needs a rubric; string matching cannot provide it
One illustration. Engine A returns one citation in each of ten answers; Engine B returns ten citations in each of ten answers. Your domain appears in five answers on each. Your reach is 50% on both. Your citation share is 50% on A and 5% on B. Neither calculation is wrong; they answer different questions. Do not announce a tenfold visibility gap without deciding which question matters.
12. Diagnose the pattern before assigning work
The purpose of measurement is to identify a solvable problem. Start with the observed combination of availability, mention, recommendation, citation and paid exposure, and walk this tree before anyone opens a content brief.
Is the response eligible and preserved? If not, fix collection, retrieval or parsing. A missing answer is not evidence that the company lost a recommendation.
Is the apparent appearance in organic prose? If it exists only in a paid unit, source-card excerpt, image target or URL, route it to the correct exposure category.
Is the right business identified? If the name is ambiguous, inspect the surrounding text, linked domain and product context before drawing a marketing conclusion.
Does the answer recommend the business for the stated use case? If the name appears in a rejection or limitation, investigate that claim. Raw mention count is not the goal.
Is the supporting information accurate and current? If not, identify the specific false or outdated assertion, its source, and the authoritative evidence that should replace it.
Is the business mentioned but its own site absent? Check whether third-party pages carry the evidence and whether your own pages answer the same specific questions clearly.
Is the business absent across repeated relevant prompts? Examine category eligibility, product fit, brand ambiguity and the candidate set before assuming a formatting change will fix it.
Does the difference survive comparable repeated measurements? If not, document it as unstable. If yes, prioritise by customer importance and tractability.
Five recurring patterns and the work they suggest
Observed pattern
First investigation
Sensible next action
Mentioned favourably, own site not cited
Which external pages support the recommendation?
Improve the authoritative page for the specific claim; assess legitimate third-party coverage
Mentioned negatively with an outdated fact
Which fact and source are wrong?
Correct your own documentation, record the correction date, use factual correction channels
Own documentation cited, no shortlist presence
Is the prompt about implementation or purchase?
Decide whether this is already a useful support outcome; examine discovery questions separately
Did surface, request, timing, parser or eligibility change?
Validate the measurement, inspect paired examples, then investigate engine-specific source changes
These actions are hypotheses to test. The study does not establish that any one change guarantees a citation.
13. What to change on your own site
The first content improvement should make a customer's important question easier to answer accurately. For a pricing question, state the plan, billing basis, currency, major limits and the date of the information. For an integration question, distinguish native support from a third-party connector or a custom API build. For a security question, distinguish a product feature, an audited certification and a contractual commitment.
A useful page puts the answer and its conditions near each other. "Unlimited users" becomes misleading if the qualifying plan, billing conditions or workspace limits live somewhere else. The same problem affects humans and machines. Use concrete evidence where it exists: reproducible measurements, dated documentation, named methodologies, release notes, worked examples and original data with clear definitions. Avoid superlatives the team cannot defend when a prospect asks for proof.
The content repair brief
Customer question: the exact question the page should answer.
Current answer problem: the specific incorrect, incomplete or missing statement.
Authoritative fact: the correct information, with its internal owner.
Evidence source: the document, test, policy or product behaviour supporting it.
Conditions: plan, geography, version, date and other limitations.
Page change: the smallest clear change that makes the fact usable.
Measurement cohort: the prompts that should plausibly be affected.
Comparison cohort: relevant prompts not targeted by the change.
Review date: when the team will assess accuracy and repeated observations.
Success definition: an answer-quality outcome, not merely a higher mention count.
This connects the content work to a falsifiable expectation. If nothing changes, the team can learn whether the page was retrieved, whether the claim was incorporated, or whether another source continues to dominate the answer.
14. A 30-day operating plan
A planning template, not a promise that AI results change within thirty days.
Days
Work
Deliverable
1 to 5
Select the customer segments and decisions that matter commercially. Gather actual questions and label generated variations separately. Record your domain inventory, aliases, competitors and product-fit boundaries. Agree the outcome rubric before looking at any answer.
A prompt and entity registry with an owner, scope and measurement rules
6 to 10
Capture the selected surfaces under documented settings. Preserve raw responses, task IDs, request configuration, timestamps and parser version. Repeat a predefined subset. Inspect a sample of wins, losses and zero-result cases.
A baseline with outcome counts, reviewed examples and a list of measurement limitations
11 to 17
Prioritise issues that combine customer importance, clear factual evidence and a change the team controls. One owner per change. Record the content before editing and the date the new information went live. Keep a comparison group of prompts the change is not meant to affect.
A short intervention log with explicit expected outcomes
18 to 24
Re-run the relevant prompts with the original settings. Use repeated captures to understand variability. Compare the same prompt-engine pairs, not the most flattering examples from each date.
A paired review with gains, losses, unchanged cases and unresolved observations
25 to 30
Continue changes that improved answer accuracy or useful exposure under the defined measurement. Revisit unclear results. Stop treating a tactic as established when the evidence remains weak.
An operating review a sceptical colleague can reproduce
A minimum weekly scorecard
Show requested and eligible answers first. Then report brand mentions, recommendation quality, own-site reach, important factual errors and paid exposure separately, each with the matched change from the previous comparable period and the number of observations behind it. "Own-site reach fell because three integration prompts now cite an outdated third-party page" is actionable. "Visibility score down 7%" is incomplete without a definition, denominator and explanation.
15. What to demand from a measurement provider
One result shown end to end: request, raw response, parsed fields, displayed metric. The demonstration should include an advertisement, a no-answer outcome, a repeated capture and a corrected historical parse.
Which surface it measures, how it identifies your company, how it treats citations and shopping links, and whether a retry creates a new observation.
Capture time distinguished from fetch time and classification time.
Historical results regenerable from preserved evidence. A parser correction should update the interpretation of the same observation, with a recorded version, not silently replace history.
A review process for false positives and false negatives. An ambiguous brand name needs a policy. A recommendation metric needs a rubric. A proprietary aggregate score needs enough explanation to support a business decision.
An example where the provider's own research claim was corrected. A system that measures changing external products needs a correction process, and this report contains several of ours.
16. Limitations and what is still open
Provider generation and completion timestamps do not exist in the study schema. Every elapsed time is a local submission, retrieval or storage time, and no result establishes when an answer was generated or whether it was served from a cache.
The longitudinal evidence runs on stored historical parser outputs and has not yet been reparsed through the corrected shared core services. The evidence file carries that warning in its own header.
The retry-sensitivity check recomputes domain Jaccard and brand flip, not retention. Retention has not been recomputed under retry exclusion.
July's task ledger shows 83.19% of completed tasks with more than one submission, a pattern consistent with a bulk resubmission rather than targeted retries, so July's retry policy cannot be established from that field.
No side-by-side interface test was run for AI Mode, so its source contraction cannot be attributed or ruled out the way ChatGPT's can.
The geographic probe is one city, one state, one IP and a ten-prompt positive control. The direct chatgpt.com ad check is one IP, one browser and three prompts.
Recomputing the original July write-ups against the frozen snapshot showed that most of their hand-typed figures do not reproduce exactly under current dedup and eligibility rules (17 of 19 page-anatomy claims, 8 of 9 video claims). Every number in this report is the recomputed value.
The prompt sample is not representative of all AI users, and nothing here measures user demand, platform trust, hidden retrieval steps, or the return on a marketing intervention.
17. Evidence register and reproducibility
Every figure in this report is computed from one frozen snapshot of the study database (SHA-256 019dbe8f1b9f8bf484ff71fefa179a9717f6055dceddf3a0235b0ee29208cb7c) and names the script that regenerates it. Superseded evidence directories computed against earlier snapshots are retained for the audit trail and are not quoted.
Cite as: Jorge Ferreiro (2026). AEO, GEO and SEO Research 2026: The AI Visibility Report. OMNIFAMOUS. https://omnifamous.com/aeo-research. Quotations are welcome with attribution and a link; please keep the cohort and denominator attached to any number you quote.
Take the report with you.
The PDF is the same document with a cover and a table of contents. The AI prompt hands the whole report to ChatGPT, Claude or Perplexity and asks it to apply the findings to your site.