AEO, GEO and SEO research in 2026. What AI actually cites.

The AI Visibility Report: what ChatGPT, Perplexity, Gemini, Google AI Mode and AI Overviews actually cite, measured on 7,949 prompts across three waves.

7,949
prompts, 158 categories
5
AI engines measured
3
collection waves, July to September
39,572
eligible answers in the final wave
239,841
citations, 27,765 domains
1.18M
citation rows in the archive

Research edition: July, August and September 2026 · OMNIFAMOUS Research

AEO, GEO and SEO Research 2026

The AI Visibility Report

Jorge Ferreiro, OMNIFAMOUS. Published 2026-09-06. Every figure names its cohort and reproduces from one frozen snapshot.

7,949
prompts, 158 categories
5
AI engines measured
3
collection waves, July to September
39,572
eligible answers in the final wave
239,841
citations, 27,765 domains
1.18M
citation rows in the archive

Executive summary

A prospect asks an AI assistant which product to buy. Your company appears in the answer, but the link goes to a comparison site. A second prospect asks a similar question and sees a competitor. Your team checks again, gets a third answer, and starts arguing about whether the content strategy is working. This report exists so that argument can be settled with evidence instead of screenshots.

Between 29 July and 5 September 2026 we submitted the same 7,949 prompts, drawn from 158 categories and 29 question archetypes, to five AI answer surfaces: ChatGPT, Perplexity, Gemini, Google AI Mode and Google AI Overview. We did it three times, roughly three weeks apart, and beside two of the waves we ran a 1,000-prompt repeat panel so that every cross-wave change could be compared with how much each engine disagrees with itself on the same day. The final wave closed with 39,572 eligible answers, 239,841 citations and 27,765 distinct cited domains. The archive behind the study holds 125,863 answer rows and 1,181,828 citation rows.

Take the full report with you, or keep reading below.

Key findings

Ten numbers, each with its caveat attached

Every figure below clears the engine's own same-day noise floor or says why it cannot. Nothing here is a causal claim.

01
14.8% to 64.6%

of July's cited domains survived to September

AI Mode kept 14.8%, ChatGPT 20.4%, Gemini 29.5%, AI Overview 42.3%, Perplexity 64.6%. Every figure clears that engine's own same-day noise floor.

02
5.40 from 15.51

sourced pages per ChatGPT answer, August to September

AI Mode fell from 17.31 to 3.34 on the same prompts. Perplexity and Gemini barely moved. A vendor interface change sits inside the ChatGPT interval.

03
44.9% to 2.7%

of ChatGPT answers citing Reddit, July to September

AI Mode fell from 60.2% to 21.5%. Gemini did not move. AI Overview rose, then fell below July. A two-point trend predicted the third point wrong twice.

04
80.0% to 0.0%

of ChatGPT answers carrying an ad, August to September

Confirmed by two independent tests. The surface that is growing is Google AI Mode: 0.21% of answers in July, 8.94% in September.

05
64.1%

of missing AI Overviews appeared on a same-day retry

237 of 370 empty captures returned an AI Overview minutes later. A once-a-day tracker reports two thirds of those absences as fact.

06
42.07%

of AI citations point to brand-owned domains

Review publishers take 17.60%, social video 7.32%, forums 5.31%, news 2.03%. 57.72% of cited pages were not on Google's page one for the same query.

07
5.61% to 34.66%

of Google page-one results an engine also cites

ChatGPT cites 5.61% of page one, AI Overview 34.66%. On ChatGPT, 70.81% of citations go to a domain that was not on page one at all.

08
55.6% vs 25.5%

page-one written pages vs videos cited by any engine

Written pages yield 1.9 answers per asset, YouTube videos 1.37. ChatGPT and Perplexity nearly stopped citing YouTube; Google surfaces did not.

09
9 of 10

local-intent answers relocated by a city target on AI Mode

Against 1 and 2 of 10 in the control arms. On brand and product prompts, geography changed nothing: 0 of 30 in every arm on every engine.

10
9.76% or 38.25%

brand flip rate on ChatGPT, depending on the denominator

The same 8,429 transitions divided by 86,393 opportunities or by the 22,037 where the brand ever appeared. Neither is a customer's risk of losing a mention.

Executive summary

A prospect asks an AI assistant which product to buy. Your company appears in the answer, but the link goes to a comparison site. A second prospect asks a similar question and sees a competitor. Your team checks again, gets a third answer, and starts arguing about whether the content strategy is working. This report exists so that argument can be settled with evidence instead of screenshots.

Between 29 July and 5 September 2026 we submitted the same 7,949 prompts, drawn from 158 categories and 29 question archetypes, to five AI answer surfaces: ChatGPT, Perplexity, Gemini, Google AI Mode and Google AI Overview. We did it three times, roughly three weeks apart, and beside two of the waves we ran a 1,000-prompt repeat panel so that every cross-wave change could be compared with how much each engine disagrees with itself on the same day. The final wave closed with 39,572 eligible answers, 239,841 citations and 27,765 distinct cited domains. The archive behind the study holds 125,863 answer rows and 1,181,828 citation rows.

The headline result is turnover. On every one of the five surfaces, the set of domains an engine cited in July had substantially changed by September, and on every surface that change is larger than the engine's own same-day variability. AI Mode retained 14.8% of its July domains, ChatGPT 20.4%, Gemini 29.5%, AI Overview 42.3% and Perplexity 64.6%. Two surfaces also cut the number of sources they expose per answer by roughly two thirds. ChatGPT's ad unit, present in 80% of August answers, was absent from every September answer. And two thirds of the AI Overviews that were missing on first capture appeared when the same query was resubmitted minutes later.

The second result is about ownership. Brand-owned domains are the single largest class of cited source (42.07% of citations), review publishers second (17.60%), and 57.72% of cited pages were not on Google's first page for the same query. What AI cites and what Google ranks are related but different inventories, and the difference is largest on ChatGPT, where 70.81% of citations go to a domain that was not on page one at all.

The third result is methodological, and it is the one we would ask you to take away if you take away only one thing. A single capture of an AI answer is an observation, not a ranking. Retention, reach, share and availability each need a stated denominator and a stated retry policy, and the same brand transitions can be reported as 9.76% or 38.25% depending on which one you pick. Chapters 11 to 14 turn that into a measurement worksheet, a diagnostic tree, a content repair brief and a 30-day operating plan you can run on your own categories.

1. What we measured and how

An engine name is shorthand for a measurement surface. Every capture in this study recorded the provider, interface, country (US), language (English), request settings and, where the provider supplied it, the serving model. Distinct provider task IDs establish separate requests; they do not establish independent generation or the absence of provider caching, and this report never assumes either.

Eligibility requires a completed task and a non-empty answer body. Copilot is excluded from every cross-engine comparison because its prompt coverage is incomplete (954 July requests only). Retries of the same provider task are not new observations. A failed task is missing evidence, not a zero-brand answer.

Collection outcomes by run
RunRequested keysEligible answersAI Overview eligibleEvery other surface
study-2026-07-2939,74539,5467,765 / 7,949 (97.68%)7,937 to 7,949 of 7,949
study-2026-08-1933,00032,4926,112 / 6,600 (92.61%)6,580 to 6,600 of 6,600
panel-2026-08-195,0004,933933 / 1,000 (93.30%)1,000 of 1,000
study-2026-09-0539,74539,5727,776 / 7,949 (97.82%)7,949 of 7,949
panel-2026-09-055,0004,968968 / 1,000 (96.80%)1,000 of 1,000

Requested keys are distinct run-prompt-engine identities. The September AI Overview figure includes 196 answers obtained by a same-day resubmission; excluding them, September AI Overview eligibility is 7,580 of 7,949 (95.36%).

The analytic samples behind each comparison
AnalysisQuestion it answersSize
Full September descriptionWhat did the engines return across the whole prompt set?39,572 eligible answers
Three-wave complete-case comparisonWhat changed on the same prompt-engine keys in all three waves?32,326 keys
Expanded July-to-September comparisonWhat changed over the longer interval, including prompts outside the August cohort?39,405 eligible pairs
Repeated September panelHow much do two same-day captures of the same prompt differ?4,958 eligible pairs

These cohorts have different sizes and must not be interchanged in a headline. "We collected nearly forty thousand requests" does not mean a given comparison contains forty thousand matched pairs.

The prompt set was built from customer language rather than from keyword tools: sales notes, support questions, on-site search and interviews, with synthetic variations labelled separately. Each prompt is tagged with the decision it represents (category discovery, fit, comparison, pricing, objection, integration, switching, implementation). A prompt portfolio has a purpose; it is not a pile of phrases a company happens to perform well on.

2. The cited web turned over in four months

This is the result the study was collected to produce. On every surface, the set of domains an engine cited in July had substantially changed by September, and on every surface the change is larger than that engine's own same-day disagreement with itself.

Cited-domain retention, July to later waves, with each engine's same-day floor
SurfaceJuly to August (pairs)July to September (pairs)Same-day floorClears the floor
ChatGPT27.7% (6,529)20.4% (7,865)52.0%yes
Perplexity71.0% (6,511)64.6% (7,847)94.0%yes
Gemini29.3% (6,189)29.5% (7,495)43.3%yes
Google AI Mode32.5% (6,567)14.8% (7,935)41.3%yes
Google AI Overview49.4% (5,918)42.3% (7,539)63.1%yes

Cohorts: July to August is the matched main-wave cohort (32,361 eligible pairs over 6,600 prompts). July to September is the expanded cohort (39,405 eligible pairs over all 7,949 prompts). Retention is the share of a July capture's cited domains that reappear in the later capture, averaged over answer pairs; pairs with an empty July set are excluded, which is why pair counts differ.

Like-for-like series on the three-wave complete-case cohort
SurfacePairsJuly to AugustJuly to SeptemberAugust to September
ChatGPT6,59027.7%20.8%37.0%
Perplexity6,60071.0%64.6%74.5%
Gemini6,59829.3%29.7%33.6%
Google AI Mode6,57932.5%15.0%15.4%
Google AI Overview5,95949.4%42.9%44.7%

Two engine-specific readings survive. Gemini's July-to-September retention (29.7%) is indistinguishable from its July-to-August retention (29.3%): whatever turnover happened, happened by August and did not compound. AI Mode is the opposite: its August-to-September retention (15.4%) is barely above its July-to-September retention (15.0%), so almost all of AI Mode's turnover occurred in the second, shorter interval.

How the noise-floor gate works

Retention and Jaccard are similarity measures. A cross-wave value below the floor means the two waves agree with each other less than the engine agrees with itself on the same day, which is exactly what makes the change real rather than noise. The floor used is the August panel, which is the stricter of the two available: September's panel agreed with itself more on every surface, so gating against August makes the finding harder to obtain, not easier.

Same-day self-agreement, broad citation policy (denominators in brackets)
SurfaceAugust panel retentionSeptember panel retentionAugust JaccardSeptember Jaccard
Perplexity94.0% (1,000)98.0% (1,000)0.911 (1,000)0.975 (1,000)
Google AI Overview63.1% (898)70.0% (919)0.474 (902)0.552 (927)
ChatGPT52.0% (998)53.9% (996)0.331 (1,000)0.398 (998)
Gemini43.3% (774)50.0% (896)0.280 (884)0.349 (953)
Google AI Mode41.3% (961)53.0% (882)0.264 (996)0.370 (948)

The September panel was drained together with the main wave (median fetch gap 0.000 hours, p90 at most 0.012 hours). That removes the August panel's ordering problem and does nothing about caching: two tasks retrieved within the same second are exactly the case where a provider cache would be least visible.

3. Two engines now expose far fewer sources

Retention fell while the number of exposed sources fell with it. On the three-wave complete-case cohort, the mean count of distinct sourced pages per answer moved as follows.

Distinct sourced pages per answer, three-wave cohort
SurfacePairsJulyAugustSeptember
ChatGPT6,59012.7015.515.40
Perplexity6,6009.8710.0310.06
Gemini6,5983.272.472.79
Google AI Mode6,57917.4917.313.34
Google AI Overview5,9599.6410.077.31

Perplexity and Gemini barely moved. ChatGPT and AI Mode, the two surfaces with the largest retention falls, are also the two that ended September exposing roughly a fifth to a third as many sourced pages per answer as in August. "The engines changed which sources they cite" and "the engines expose fewer sources" are both consistent with these numbers, and this corpus cannot separate them. A retention figure computed against a shrinking target set will fall even if the surviving choices are unchanged in kind.

For ChatGPT there is a documented capture-side explanation. The collection vendor's changelog records that between 25 and 26 August its mobile-web format began serving 98% of ChatGPT responses, with metadata fields potentially empty, and that from 28 August a legacy flag requests the older desktop interface on a best-effort basis. Our own side-by-side test on 5 September (20 prompts per arm) found the current default recovered a mean of 0.1 fan-out search queries per answer against 1.8 on the legacy desktop arm, with a different response key set and a different mean answer length (3,690 against 2,688 characters).

4. Reddit and YouTube: a two-point trend was wrong twice

This is the strongest available argument against building a content programme on one wave's source leaderboard. Reach counts an eligible answer once when the domain appears in its sourced evidence. All three columns are the same prompt-engine keys.

Reddit reach by wave, three-wave cohort
SurfacePairsJulyAugustSeptemberJuly to September change [95% CI]
ChatGPT6,59044.90%8.01%2.69%-42.22 pp [-43.43, -41.00]
Perplexity6,60027.58%23.65%22.71%-4.86 pp [-5.71, -4.02]
Gemini6,5988.99%8.84%9.05%+0.06 pp [-0.75, +0.87]
Google AI Mode6,57960.24%42.48%21.46%-38.77 pp [-40.13, -37.42]
Google AI Overview5,95949.94%52.53%39.47%-10.47 pp [-11.91, -9.04]
YouTube reach by wave, three-wave cohort
SurfacePairsJulyAugustSeptemberJuly to September change [95% CI]
ChatGPT6,5903.02%4.66%0.20%-2.82 pp [-3.24, -2.40]
Perplexity6,60028.88%1.65%0.59%-28.29 pp [-29.38, -27.20]
Gemini6,5983.06%2.18%2.23%-0.83 pp [-1.34, -0.33]
Google AI Mode6,57956.70%58.46%40.14%-16.55 pp [-18.00, -15.10]
Google AI Overview5,95958.73%60.18%55.48%-3.26 pp [-4.61, -1.91]

Intervals use a prompt-cluster, ratio-of-sums normal approximation. They do not adjust for category-level shocks, multiple comparisons, measurement error or changes to the provider surface.

Four readings, in decreasing order of confidence. Perplexity's YouTube reach fell by an order of magnitude in the first interval and stayed down. Gemini did essentially nothing on either domain across four months, so "AI cites Reddit less now" is not a statement about AI. AI Overview's Reddit reach rose from July to August and then fell below its July level, so a team that had extrapolated the first interval would have been wrong about the second. And ChatGPT's falls, on both domains, are the ones to trust least, because they sit inside the documented interface change and coincide with its sourced-page count dropping from 15.51 to 5.40.

What none of this supports is an instruction. It does not show that forum participation stopped working, that video stopped working, or that either will still be down next month. It shows that the composition of the cited web on these surfaces moved a long way in four months, which is an argument for measuring your own categories repeatedly rather than for reallocating a budget on the strength of a single leaderboard.

5. ChatGPT ads went to zero. AI Mode ads are growing

Advertisements create a measurement problem because they may arrive as separate structured data, appear in rendered prose, or both. A paid slot beside an answer is not an organic recommendation, and a rendered ad can inject a competitor name into organic mention detection if the two are not separated at parse time.

Share of eligible answers carrying at least one detected ad, three-wave cohort
SurfacePairsJulyAugustSeptemberAugust to September change [95% CI]
ChatGPT6,59069.74%80.02%0.00%-80.02 pp [-80.98, -79.05]
Google AI Mode6,5790.21%4.38%8.94%+4.56 pp [+3.79, +5.33]
Google AI Overview5,9590.25%0.49%0.35%-0.13 pp [-0.36, +0.09]
Perplexity6,6000.00%0.00%0.00%+0.00 pp
Gemini6,5980.00%0.00%0.00%+0.00 pp

The ChatGPT zero is not a parsing gap and not a rounding artefact. The structured ad table holds 10,380 ChatGPT rows for August and zero rows of any provenance for September, while AI Mode and AI Overview ad rows kept flowing in the same September run. Sampling 200 raw payloads per wave, 171 of 200 August payloads (85.5%) carried a non-empty ads array against 0 of 200 in September.

Two independent tests rule out the two obvious measurement explanations. The vendor's older desktop interface, requested on the same 20 prompts on the same day, returned zero ads on 18 of 18 completed tasks. And chatgpt.com itself, checked directly in a logged-out browser on 5 September, served no ad markup on three commercial prompts, two of which had carried 15 ad slots each in August.

Meanwhile the surface that is quietly growing is Google AI Mode, from 0.21% of answers in July to 8.94% in September on the same 6,579 pairs. Its ad provenance is structured in all three waves and all three were parsed under the same harmonised parser, so the specific confound that destroys the ChatGPT July-to-August comparison does not apply. On the evidence available it is the paid surface worth watching next.

One correction to our own earlier work belongs here. Regenerating the July analysis against the frozen snapshot showed that its headline claim, that ChatGPT was the only engine serving this ad unit, was already false in July: 17 AI Mode answers and 22 AI Overview answers carried a structured ad in the July run, on a rounding-to-zero scale that an array-only parser returned as literal zero.

6. An absent AI Overview is often a property of the capture

The September wave closed with 383 non-completed tasks: 370 empty, every one of them AI Overview, and 13 failed. All 383 were resubmitted the same day, unchanged.

Outcome of the same-day retry on 370 empty AI Overview captures
OutcomenShare of the 370 empties
Returned an AI Overview23764.1%
Still empty13335.9%

All 13 failures also recovered. The 133 residual empties reconcile as 113 on the main wave and 20 on the panel.

Two thirds of the queries that had "no AI Overview" produced one minutes later, on the same query, same country, same day. Three consequences follow, and the third is the commercial one.

  1. A prevalence claim has to state its retry policy. "AI Overview appears on X% of prompts", measured from a single capture, understates X by roughly the share of transient empties, which here was 64% of observed empties.
  2. A cross-wave coverage comparison is confounded unless both waves retried the same way. September's AI Overview eligibility of 97.82% against August's 92.61% cannot be read as Google expanding coverage. Excluding September's retry-obtained captures brings it to 95.36%, still above August and still below July's 97.68%, which is a mixed picture rather than a trend.
  3. A once-a-day monitoring product inherits this directly. It captures once per prompt per day and reports an absent AI Overview as an absence. On this evidence roughly two in three of those absences would not survive a second look taken minutes later.

7. Who owns the cited web

Every cited domain in the September wave was classified by owner type using a layered scheme: manual overrides, curated exact-domain lists, TLD rules, host aliasing, a seed-brand match and evidence-based classification from page paths and titles. 9,119 of 27,765 domains were classified (32.84%), covering 86.77% of citation volume. The long tail is real: 18,646 unclassified domains hold 13.23% of citations.

Citations by owner type, September wave (39,572 eligible answers)
Owner typeDomainsCitationsShare of citationsAnswers reachedShare of answers
Brand-owned5,539100,89742.07%28,95073.16%
Review publisher1,49742,21117.60%18,80047.51%
Unclassified18,64631,72513.23%16,40741.46%
Social video1217,5517.32%8,93722.58%
Other1,71015,6356.52%9,57124.19%
Forum (UGC)1912,7265.31%8,28120.93%
Aggregator / directory835,9442.48%3,9209.91%
News774,8742.03%3,3628.50%
Retailer593,2351.35%1,3033.29%
Government / education / medical962,6981.12%1,4043.55%
Docs / code162,1970.92%1,1042.79%
Affiliate111480.06%1410.36%
The 20 domains reaching the most answers, September wave
RankDomainOwner typeAnswers reached
1youtube.comSocial video7,821
2reddit.comForum7,549
3g2.comReview publisher4,795
4capterra.comReview publisher1,815
5zapier.comBrand-owned1,558
6trustpilot.comReview publisher1,500
7forbes.comNews1,282
8google.comOther1,248
9medium.comOther1,158
10apple.comOther1,116
11github.comDocs / code1,008
12sourceforge.netAggregator907
13gartner.comReview publisher899
14linkedin.comSocial video856
15techradar.comReview publisher830
16softwareadvice.comReview publisher792
17facebook.comSocial video721
18pcmag.comReview publisher701
19slashdot.orgAggregator606
20amazon.comRetailer598

Reach counts an answer once per domain, regardless of how many times the domain appears in it. google.com is largely Play Store listings, apple.com mixes App Store with support, and amazon.com mixes retail with AWS; the taxonomy flags these wherever it matters. A leaderboard is a description of this prompt set, not a publishing strategy.

AI citations are less concentrated than Google's first page

Concentration, size-matched draws of 66,366 citations (30 draws)
MetricAI cited (mean, sd)Google page one (mean, sd)
Domains needed for 50% of citations533.2 (5.9)328.7 (5.0)
Top-10 domain share19.58%25.36%
Herfindahl index76.7 (0.8)143.9 (1.7)
Effective number of domains130.469.5
Concentration by engine, full depth, September wave
EngineCitationsDomainsDomains for 50%Top-10 shareGini
Google AI Mode30,6096,54511836.03%0.707
Google AI Overview60,07111,25225429.27%0.726
ChatGPT46,5169,51252211.93%0.677
Gemini22,6397,93073811.57%0.555
Perplexity80,00614,24246016.12%0.726

The Google-owned surfaces concentrate their citations on far fewer domains than ChatGPT, Gemini or Perplexity do. AI Mode needs 118 domains to account for half of its citations; Gemini needs 738. For a brand, that is the difference between a surface where a handful of publishers carry most of the evidence and a surface where the long tail is still in play.

Vendor sites are cited, but their help centres barely are

Of the 1,832 seed vendor domains in the study, 1,690 (92.25%) were cited at least once in the September wave. The median vendor domain reached 10 answers, the 90th percentile 46, and the maximum 1,570. Of the 55,859 citations to vendor domains, 2.07% pointed at a docs subdomain, 1.86% at help, 1.44% at support, 1.11% at blog and 0.60% at status. The marketing site is what gets cited; the documentation that would actually settle a fit or integration question is a rounding error.

8. What Google ranks and what AI cites are different inventories

For 7,938 of the 7,949 prompts, the study also captured Google's first page of organic results (66,366 rows, mean depth 8.4). That allows one question per engine: of the pages Google ranked on page one for this query, how many did the engine also cite?

Share of Google page-one results also cited, by engine and rank, September wave
EngineShare of page one citedCited at rank 1Cited at rank 9Rank-1 to rank-9 lift
Google AI Overview34.66%48.15%25.23%1.9x
Perplexity24.53%39.13%14.88%2.6x
Google AI Mode12.79%27.21%5.55%4.9x
Gemini11.23%15.32%6.40%2.4x
ChatGPT5.61%11.49%2.54%4.5x
Where an engine's citations come from, relative to Google page one
EngineCitationsExact page on page oneSame domain, different pageDomain not on page one
Google AI Overview60,00737.45%19.56%42.98%
Google AI Mode30,58127.77%22.25%49.98%
Gemini22,60532.94%11.15%55.91%
Perplexity79,90920.38%15.16%64.45%
ChatGPT46,4938.02%21.18%70.81%

Rank still matters inside every engine: a page at position one is 1.9 to 4.9 times more likely to be cited than a page at position nine. But the overall overlap is small everywhere except AI Overview, and on ChatGPT seven in ten citations go to a domain that Google did not rank on page one at all. Across all engines, 57.72% of cited pages were off Google's first page. The off-page-one rate is lowest for forum citations (10.40%), because Reddit ranks, and highest for the unclassified long tail (83.24%).

9. Written pages beat video as cited assets

Share of answers citing YouTube, by engine
EngineJuly answersJuly citing YouTubeSeptember answersSeptember citing YouTube
ChatGPT7,8652.8%7,8850.3%
Perplexity7,84728.9%7,9440.6%
Gemini7,4953.1%7,1752.3%
Google AI Mode7,93556.6%7,22946.1%
Google AI Overview7,67960.1%7,47757.1%

Among assets that Google ranked on page one, a written page was cited by at least one engine 55.6% of the time, a UGC or social page 36.0% of the time, and a YouTube video 25.5% of the time. Per asset, written pages yielded a mean of 1.9 answers (over 111,261 assets) against 1.37 for YouTube videos (over 10,818). Video is a Google-surface phenomenon: AI Mode and AI Overview kept citing it in roughly half of answers while ChatGPT and Perplexity all but stopped.

A cited-video sample omits the videos that were never cited, so none of this identifies the causal effect of making a video. It does say that if your category is answered mainly by ChatGPT or Perplexity, a video-first source strategy has very little evidence behind it in this corpus.

10. Geography moves local intent, not brand visibility

Every capture in the study used a national US setting. That was an assumption, so it was tested: 40 prompts, three engines, three arms (national, an identical national repeat as the same-day floor, and a San Francisco target), 360 tasks. The prompt set was 30 corpus prompts plus 10 local-intent prompts as a positive control.

Answers naming San Francisco, the Bay Area, Oakland, Berkeley or the Mission
EngineMechanismPrompt groupNationalNational repeatCity target
Google AI Modelocation: cityLocal intent2/101/109/10
ChatGPTstate: CALocal intent0/100/102/10
Perplexitystate: CALocal intent0/101/100/10
All threeeitherBrand and product prompts0/300/300/30

AI Mode's city target genuinely relocates the answer. A state-level proxy is much weaker, which is what a state-level proxy should be. Perplexity appears to ignore the state parameter entirely: its cited domains were byte-identical across all three arms on all 40 prompts. And on brand and product questions, geography did nothing at all.

11. Six outcomes that should never share one score

The most consequential measurement choice is deciding what counts as a success. An answer can mention your company without recommending it. It can recommend you without citing your website. It can cite your documentation without putting your product on a shortlist. A sponsored placement can sit beside the answer and create an apparent organic win. If those outcomes become one score, the score conceals the decision your team needs to make.

The six outcomes
OutcomeDefinitionWhat it is not
Answer availabilityThe surface returned an eligible answer for the promptWhether your brand appears
Brand mentionThe organic prose names your company or a verified aliasA name inside a tracking URL, image target, source-card excerpt or ad
RecommendationThe answer suggests your product for the stated use caseA negative comparison that happens to contain your name
Own-site citationAn eligible organic source points to a domain you ownAn endorsement; it may support one narrow fact
Inline destinationA navigable link to your business inside the proseA source listed in a separate citation module
Paid or commerce exposureAn ad, shopping destination or product cardAn organic recommendation

A brand mentioned in 40 eligible answers and cited in 10 has two measured outcomes. Reporting a combined "50 visibility events" double-counts some observations and removes the information needed to improve the result.

A denominator can make the same brand transitions look four times larger

For ChatGPT, the seeded category-brand roster creates 86,393 eligible brand opportunities across the paired July and August prompts. There are 5,072 observed losses and 3,357 gains, totalling 8,429 flips. Dividing by all 86,393 opportunities gives 9.76%. Dividing the same 8,429 flips by the 22,037 opportunities where the brand appeared at least once gives 38.25%. Neither number is the probability that a customer account loses an AI recommendation. The first weights every category-brand opportunity, including brands absent on both captures. The second conditions on ever appearing. Both measure detected name presence, not favourable recommendation.

Every rate needs a stated population
MetricNumeratorDenominatorTrap
Brand mention rateEligible answers containing the brandEligible answers in scopeReport availability separately
Domain reachEligible answers containing the domain at least onceEligible answersTen citations in one answer is one reached answer
Citation-row shareThe domain's eligible citation rowsAll eligible citation rowsSensitive to how many sources each answer carries
RetentionDomains in both capturesDomains in the first captureUndefined when the first set is empty; never insert zero
JaccardIntersectionUnionBoth-empty is its own state, not perfect stability
Recommendation rateLabelled favourable recommendationsEligible answersNeeds a rubric; string matching cannot provide it

One illustration. Engine A returns one citation in each of ten answers; Engine B returns ten citations in each of ten answers. Your domain appears in five answers on each. Your reach is 50% on both. Your citation share is 50% on A and 5% on B. Neither calculation is wrong; they answer different questions. Do not announce a tenfold visibility gap without deciding which question matters.

12. Diagnose the pattern before assigning work

The purpose of measurement is to identify a solvable problem. Start with the observed combination of availability, mention, recommendation, citation and paid exposure, and walk this tree before anyone opens a content brief.

  1. Is the response eligible and preserved? If not, fix collection, retrieval or parsing. A missing answer is not evidence that the company lost a recommendation.
  2. Is the apparent appearance in organic prose? If it exists only in a paid unit, source-card excerpt, image target or URL, route it to the correct exposure category.
  3. Is the right business identified? If the name is ambiguous, inspect the surrounding text, linked domain and product context before drawing a marketing conclusion.
  4. Does the answer recommend the business for the stated use case? If the name appears in a rejection or limitation, investigate that claim. Raw mention count is not the goal.
  5. Is the supporting information accurate and current? If not, identify the specific false or outdated assertion, its source, and the authoritative evidence that should replace it.
  6. Is the business mentioned but its own site absent? Check whether third-party pages carry the evidence and whether your own pages answer the same specific questions clearly.
  7. Is the business absent across repeated relevant prompts? Examine category eligibility, product fit, brand ambiguity and the candidate set before assuming a formatting change will fix it.
  8. Does the difference survive comparable repeated measurements? If not, document it as unstable. If yes, prioritise by customer importance and tractability.
Five recurring patterns and the work they suggest
Observed patternFirst investigationSensible next action
Mentioned favourably, own site not citedWhich external pages support the recommendation?Improve the authoritative page for the specific claim; assess legitimate third-party coverage
Mentioned negatively with an outdated factWhich fact and source are wrong?Correct your own documentation, record the correction date, use factual correction channels
Own documentation cited, no shortlist presenceIs the prompt about implementation or purchase?Decide whether this is already a useful support outcome; examine discovery questions separately
Visibility appears only in adsIs the exposure paid and correctly tagged?Evaluate paid performance independently; keep organic metrics clean
One engine changes sharplyDid surface, request, timing, parser or eligibility change?Validate the measurement, inspect paired examples, then investigate engine-specific source changes

These actions are hypotheses to test. The study does not establish that any one change guarantees a citation.

13. What to change on your own site

The first content improvement should make a customer's important question easier to answer accurately. For a pricing question, state the plan, billing basis, currency, major limits and the date of the information. For an integration question, distinguish native support from a third-party connector or a custom API build. For a security question, distinguish a product feature, an audited certification and a contractual commitment.

A useful page puts the answer and its conditions near each other. "Unlimited users" becomes misleading if the qualifying plan, billing conditions or workspace limits live somewhere else. The same problem affects humans and machines. Use concrete evidence where it exists: reproducible measurements, dated documentation, named methodologies, release notes, worked examples and original data with clear definitions. Avoid superlatives the team cannot defend when a prospect asks for proof.

The content repair brief

  • Customer question: the exact question the page should answer.
  • Current answer problem: the specific incorrect, incomplete or missing statement.
  • Authoritative fact: the correct information, with its internal owner.
  • Evidence source: the document, test, policy or product behaviour supporting it.
  • Conditions: plan, geography, version, date and other limitations.
  • Page change: the smallest clear change that makes the fact usable.
  • Measurement cohort: the prompts that should plausibly be affected.
  • Comparison cohort: relevant prompts not targeted by the change.
  • Review date: when the team will assess accuracy and repeated observations.
  • Success definition: an answer-quality outcome, not merely a higher mention count.

This connects the content work to a falsifiable expectation. If nothing changes, the team can learn whether the page was retrieved, whether the claim was incorporated, or whether another source continues to dominate the answer.

14. A 30-day operating plan

A planning template, not a promise that AI results change within thirty days.

DaysWorkDeliverable
1 to 5Select the customer segments and decisions that matter commercially. Gather actual questions and label generated variations separately. Record your domain inventory, aliases, competitors and product-fit boundaries. Agree the outcome rubric before looking at any answer.A prompt and entity registry with an owner, scope and measurement rules
6 to 10Capture the selected surfaces under documented settings. Preserve raw responses, task IDs, request configuration, timestamps and parser version. Repeat a predefined subset. Inspect a sample of wins, losses and zero-result cases.A baseline with outcome counts, reviewed examples and a list of measurement limitations
11 to 17Prioritise issues that combine customer importance, clear factual evidence and a change the team controls. One owner per change. Record the content before editing and the date the new information went live. Keep a comparison group of prompts the change is not meant to affect.A short intervention log with explicit expected outcomes
18 to 24Re-run the relevant prompts with the original settings. Use repeated captures to understand variability. Compare the same prompt-engine pairs, not the most flattering examples from each date.A paired review with gains, losses, unchanged cases and unresolved observations
25 to 30Continue changes that improved answer accuracy or useful exposure under the defined measurement. Revisit unclear results. Stop treating a tactic as established when the evidence remains weak.An operating review a sceptical colleague can reproduce

A minimum weekly scorecard

Show requested and eligible answers first. Then report brand mentions, recommendation quality, own-site reach, important factual errors and paid exposure separately, each with the matched change from the previous comparable period and the number of observations behind it. "Own-site reach fell because three integration prompts now cite an outdated third-party page" is actionable. "Visibility score down 7%" is incomplete without a definition, denominator and explanation.

15. What to demand from a measurement provider

  • One result shown end to end: request, raw response, parsed fields, displayed metric. The demonstration should include an advertisement, a no-answer outcome, a repeated capture and a corrected historical parse.
  • Which surface it measures, how it identifies your company, how it treats citations and shopping links, and whether a retry creates a new observation.
  • Capture time distinguished from fetch time and classification time.
  • Historical results regenerable from preserved evidence. A parser correction should update the interpretation of the same observation, with a recorded version, not silently replace history.
  • A review process for false positives and false negatives. An ambiguous brand name needs a policy. A recommendation metric needs a rubric. A proprietary aggregate score needs enough explanation to support a business decision.
  • An example where the provider's own research claim was corrected. A system that measures changing external products needs a correction process, and this report contains several of ours.

16. Limitations and what is still open

  • Provider generation and completion timestamps do not exist in the study schema. Every elapsed time is a local submission, retrieval or storage time, and no result establishes when an answer was generated or whether it was served from a cache.
  • The longitudinal evidence runs on stored historical parser outputs and has not yet been reparsed through the corrected shared core services. The evidence file carries that warning in its own header.
  • The retry-sensitivity check recomputes domain Jaccard and brand flip, not retention. Retention has not been recomputed under retry exclusion.
  • July's task ledger shows 83.19% of completed tasks with more than one submission, a pattern consistent with a bulk resubmission rather than targeted retries, so July's retry policy cannot be established from that field.
  • No side-by-side interface test was run for AI Mode, so its source contraction cannot be attributed or ruled out the way ChatGPT's can.
  • The geographic probe is one city, one state, one IP and a ten-prompt positive control. The direct chatgpt.com ad check is one IP, one browser and three prompts.
  • Recomputing the original July write-ups against the frozen snapshot showed that most of their hand-typed figures do not reproduce exactly under current dedup and eligibility rules (17 of 19 page-anatomy claims, 8 of 9 video claims). Every number in this report is the recomputed value.
  • The prompt sample is not representative of all AI users, and nothing here measures user demand, platform trust, hidden retrieval steps, or the return on a marketing intervention.

17. Evidence register and reproducibility

Every figure in this report is computed from one frozen snapshot of the study database (SHA-256 019dbe8f1b9f8bf484ff71fefa179a9717f6055dceddf3a0235b0ee29208cb7c) and names the script that regenerates it. Superseded evidence directories computed against earlier snapshots are retained for the audit trail and are not quoted.

Claim areaEvidence fileGenerating script
Collection outcomes, retention, Jaccard, brand transitions, reach, ad rates, noise-floor gates, retry sensitivityaudits/2026-09-05/longitudinal-v5-retry-sensitivity/LONGITUDINAL_EVIDENCE.md and its JSONanalysis/longitudinal.py
Same-day repeatability floorsanalysis/NOISE_FLOOR.mdanalysis/noise-floor.mjs
September collection contract, retry outcome, availability caveatanalysis/SEPTEMBER_METHODOLOGY.mdtasks table on the snapshot
ChatGPT ad collapse, the two independent confirmationsexperiments/chatgpt-ads-legacy/FINDINGS.mdexperiments/chatgpt-ads-legacy/run.mjs
Ad rates by run and provenance, July reconciliationaudits/2026-09-05-final/chatgpt-ads/chatgpt_ads.mdanalysis/chatgpt_ads.py
Geographic targeting probeexperiments/geo-targeting/FINDINGS.mdexperiments/geo-targeting/run.mjs
Who owns the cited web, three wavesaudits/2026-09-05-final/who-owns-the-cited-web/WHO_OWNS_THE_CITED_WEB.mdanalysis/who_owns_the_cited_web.py
SERP position versus citation, three wavesaudits/2026-09-05-final/serp-vs-cited/SERP_VS_CITED.mdanalysis/serp_vs_cited.py
Page-level anatomy of cited pagesaudits/2026-09-05-final/page-anatomy/page_anatomy.mdanalysis/page_anatomy.py
Video versus written, three wavesaudits/2026-09-05-final/video-vs-written/video_vs_written.mdanalysis/video_vs_written.py
Archive storage totalsdata/db-backup-manifest.jsonbackup job

Cite as: Jorge Ferreiro (2026). AEO, GEO and SEO Research 2026: The AI Visibility Report. OMNIFAMOUS. https://omnifamous.com/aeo-research. Quotations are welcome with attribution and a link; please keep the cohort and denominator attached to any number you quote.

Take the report with you.

The PDF is the same document with a cover and a table of contents. The AI prompt hands the whole report to ChatGPT, Claude or Perplexity and asks it to apply the findings to your site.

Want the same measurement on your own categories? Run a free AI visibility audit

Get the report as a PDF

Tell us where to send it and the download starts right away. We use your site to tailor what we send you next, nothing else.

No spam. One email with the report, then only research updates.

Join the OMNIFAMOUS waitlist

Tell us about your site and we'll let you in when your spot opens up.