Your brand is described by pages you do not own

Ask an assistant which brand to buy in a category and look at what it cites. A minority of it will be the brands’ own sites.

A 2026 comparison classified cited sources as earned, social or brand-owned across consumer ranking queries. For Claude 4.5 Sonnet the mix was 65% earned and 1% social; for GPT-4o, 57% earned and 8% social; Perplexity Sonar Pro leaned harder on brand-owned pages at 39%; Gemini 2.5 Flash split roughly evenly between earned and brand (Chen et al., arXiv:2601.16858). The mix varies by engine, and in most of them earned coverage dominates.

Vendor datasets on Google’s AI Overviews point the same way, hard. Ahrefs reported in September 2026 that the most-cited domains across more than three million US queries were youtube.com at 22.9%, reddit.com at 18.5%, facebook.com at 10.1%, with Wikipedia at 4.0% (Ahrefs). Treat the exact shares as one vendor’s window, not a constant: Semrush’s analysis of 230,000 prompts recorded ChatGPT’s Reddit citation rate falling from roughly 60% of responses to about 10% within a month in late 2025 (Semrush).

The direction survives the disagreement. User-generated content, video and reference sites are consistently over-represented, and your product page is one source among several.

Why this is structural, not a phase

Two mechanisms reinforce it.

First, an assistant has room for about four sources: one large study found a mean of 4.3 URLs and 3.4 domains cited per answer, against 10.3 and 7.3 for traditional engines (Zhang et al., arXiv:2512.09483). With four slots, a single page that compares ten brands is worth more to the model than ten brand pages that each praise one.

Second, some of what the model says never came from a retrieved page at all. Models “are heavily influenced by their latent knowledge of product names” (Pfrommer et al., arXiv:2406.03589), and 16% of entities in generated rankings in the 2026 study appeared in no retrieved snippet. That prior was formed by the same third-party corpus, at training time.

What follows for the work

  • Audit the cited sources, not just your rank. For every answer in your category, record which pages were cited. That list is your real competitive set, and it belongs in your displacement metric.
  • Fix the wrong facts where they live. An outdated spec on a retailer page or a stale review is a correctable input. This is unglamorous, and it is the highest-leverage work available.
  • Earn comparison coverage. The pages assistants like are the ones that compare, test and quantify. You cannot write them about yourself credibly, but you can supply the data they need.
  • Keep the search crawlers in. Third-party pages only help if the engine can fetch them and yours, which makes the crawler decision a visibility decision, not just a legal one.
  • Expect the answer to lag. Training-time priors update slowly, which is why a burst of activity on your own domain often moves nothing, and why distinguishing retrieval from memory matters before you spend anything.

There is a commercial reason this is worth the trouble. Pew found that when an AI summary appeared, users clicked a result in 8% of visits against 15% without one, and clicked a link inside the summary in 1% of visits (Pew Research Center). The visit is increasingly not the outcome. Being described accurately is, and the description is assembled from sources you can find but not edit.