Most of the sources AI engines cite are not brand websites, and the margin is not close. In Muck Rack's May 2026 edition of What Is AI Reading?, an analysis of more than 25 million links cited by ChatGPT, Claude and Gemini across seventeen industries, earned media accounted for 84% of all citations, and paid or advertorial content for 0.3%. The figure has barely moved: across the study's three editions since July 2025, earned media has held between 82% and 89%. A narrower cut makes the same point more sharply. DerivateX's June 2026 study put B2B vendor-discovery questions to ChatGPT across forty software categories and found the engine citing a recommended product's own website just 11.6% of the time; 87.4% of the citations attached to its recommendations pointed at third-party pages. The engine names you, and hands the microphone to someone else.
July's work on this site was about the half you own: whether a crawler can reach your pages, whether your substance survives the fetch, whether a passage can be lifted, attributed and trusted. Nothing below revokes a line of it, because an unfetchable page is invisible on every surface at once. But the floor is not the ceiling. A brand can pass all five owned checks and still barely appear in answers, because the engine composing those answers is reading, mostly, what other people wrote. This piece maps where that writing lives, what each kind of room contributes, and how to find your own earned gaps in one afternoon.
Why the balance tilts earned
An answer engine has a trust problem a search engine never quite had: it does not present a list for you to judge; it asserts. Before asserting, it behaves like a careful editor, preferring claims it can find made independently in more than one place, by parties without an obvious stake. A brand's own page is, by construction, the party with the stake. This is not a penalty and it is not personal; it is what any of us does when a company's homepage says it is the best and a forum thread, a review profile and a journalist all say something more specific.
The evidence that the preference is structural rather than stylistic is the cleanest experiment in the field so far. In a controlled study run by Stacker and Scrunch in late 2025, eight articles were measured across 944 prompt-and-platform combinations on five engines: first on a brand's own site, then distributed through third-party news outlets. On the brand's domain, the content was cited in 8% of runs. Distributed, the identical content reached 34%; a broader follow-up in March 2026, covering 87 stories from thirty brands, found a 239% median lift. Eight articles is a small base, and the honest reading treats the exact multiplier as provisional. The direction, though, is not in doubt, and the direction is the finding: the room is doing work the writing cannot do alone.
One more strand ties the earned half to the classical one, and it deserves its precise credit. Lily Ray's February 2026 study of site sections demoted in that January's Google update found losses in organic visibility mirrored by declines in AI citations in nearly every engine she measured, with Perplexity the outlier that grew against the trend. The reading that matters here: the authority classical search spent two decades learning to measure, which was always mostly earned and off-site, is substantially the same authority the answer engines now consult. Earned presence is not a new tax. It is the old currency, spending in a new economy.
The earned surface map: five rooms, and what each contributes
Journalism. Roughly 27% of all citations in the May 2026 Muck Rack data are journalistic sources, a share that has stayed between 25% and 27% across every edition of the study. Journalism dominates time-sensitive and industry-trend questions: when the query implies "what is happening", engines reach for outlets. The contribution is legitimacy plus recency; the cost is that it is the slowest room to enter, and the least controllable once entered.
Communities. Forums, and Reddit above all, are where engines find unsponsored experience: the phrasing buyers actually use, the complaints vendors never publish, the recommendation made by someone with nothing to sell. Engines weight this room more heavily than most brands expect: in SE Ranking's 100,000-keyword study of Google AI Overviews, Reddit and Quora sit inside the top ten most-cited domains, beside YouTube at the top of the list and ahead of almost every traditional publisher. The contribution is authenticity; the cost is that the room has an immune system, and next week's essay is about exactly that.
Review platforms. Seer Interactive's May 2026 study of 800,000 AI responses, conducted for Trustpilot and worth reading with that commission in mind, found review and trust sites the second-highest citation source overall, with the engines quoting ratings and reviewer language directly into answers, and treating an absent or empty profile as its own signal. For commercial queries the room functions as an inclusion gate more than a ranking lever: a missing profile can keep a brand out of the comparison entirely.
Comparison and listicle pages. The pages that answer "best X for Y" are cited out of proportion to their prestige because they are shaped exactly the way retrieval wants: one question, a ranked or criterion-scored answer near the top, entities named densely, verdicts quotable in isolation. A brand page structurally cannot be this page for its own category, because the party with the stake cannot rank itself credibly. The move is not to imitate the form but to be present, and accurately represented, on the instances of it the engines already trust.
Practitioner blogs and independent analysis. The smallest room by volume and often the highest in trust per citation: a named specialist, writing under their own reputation, with method shown. Engines treat these sources the way careful readers do. The contribution is depth of trust; the entry price is being genuinely useful to the practitioner, which cannot be faked and cannot be bought.
The earned-gap analysis: one afternoon, one list
The map becomes a to-do list through a procedure any team can run without a vendor. First, take the fifteen to twenty buyer questions from your measurement baseline and run them again across the engines you care about, but this time record the sources, not just the verdicts: every domain cited in every answer. Second, classify each cited source into the five rooms above, and mark for each whether it mentions your brand, mentions competitors but not you, or mentions neither. Third, sort the "competitors but not you" list by how often each source appears across your query set. That ranked list is the earned to-do: each row is a specific page or profile the engines already trust for your exact questions, currently answering them without you in the frame. It is the rare artefact in marketing that is simultaneously a diagnosis, a priority order and an outreach list, and it costs one afternoon.
Two honesty notes before the month builds on this. The 11.6% own-site figure is one engine on one query class, vendor discovery in B2B software, where the third-party skew is at its strongest; informational questions about a brand people already know cite owned pages far more readily, and any team should expect its own split to vary by intent class. And nothing in the earned half argues against the owned work: the distributed articles in the Stacker experiment were still well-structured articles. The floor makes the ceiling reachable. July built the floor. This month is about the rooms above it.
Next week goes inside those rooms: what participation looks like when the room cannot be bought, and why showing up to give is the only strategy that survives contact with a community. The week after, the mechanism underneath all of it: how engines cross-check before they cite, and the audit that measures whether your claims corroborate.
You can own your answers. You cannot own a citation. The rest of this month is about deserving one.
Frequently asked questions
What proportion of AI citations come from earned media rather than brand websites?
Earned media accounted for 84% of all citations in Muck Rack's May 2026 analysis of more than 25 million links across seventeen industries, with paid or advertorial content at 0.3%. The share has held between 82% and 89% since July 2025. In B2B vendor discovery the skew is sharper still: DerivateX found ChatGPT citing a recommended product's own website just 11.6% of the time.
Does the same content get cited more often when published off-site?
Yes, substantially. Stacker and Scrunch measured eight articles across 944 prompt-and-platform combinations on five engines. On the brand's own domain the content was cited in 8% of runs; distributed through third-party outlets, the identical content reached 34%. A broader follow-up across 87 stories found a 239% median lift. The base is small, so treat the multiplier as provisional and the direction as settled.
Why do AI engines prefer third-party sources over a brand's own website?
Because an answer engine asserts rather than presenting a list to judge. Before asserting it behaves like a careful editor, preferring claims it can find made independently in more than one place, by parties without an obvious stake. A brand's own page is, by construction, the party with the stake. This is structural, not a penalty.
Which types of third-party sources do AI engines cite most?
Five rooms carry the earned half: journalism, at roughly 27% of citations and dominant on trend questions; communities, where Reddit and Quora sit inside the top ten most-cited domains; review platforms, which act as an inclusion gate for commercial queries; comparison and listicle pages, cited out of proportion to their prestige because their structure suits retrieval; and practitioner blogs, the smallest room by volume and often the highest in trust per citation.
How do you find your own earned citation gaps?
Run the fifteen to twenty buyer questions from your measurement baseline again, recording every domain cited in every answer rather than just the verdicts. Classify each source into the five rooms and mark whether it mentions you, mentions competitors but not you, or mentions neither. Sort the "competitors but not you" list by how often each source recurs. That ranked list is a diagnosis, a priority order and an outreach list at once.