· · GEOAI OverviewsAI Citations

Out-competed by arxiv.org in AI Answers: What Marketing Teams Do When the Citation Isn't a Competitor Brand

Your organic traffic is declining and the AI is citing arxiv.org, Wikipedia, or Reddit instead of your site. Why this pattern is different from losing to a competitor brand, how to diagnose it across ChatGPT, Claude, Google AI Overviews, and Perplexity, and the three plays marketing teams run in response.

Out-competed by arxiv.org in AI Answers: What Marketing Teams Do When the Citation Isn't a Competitor Brand

You have confirmed the drop. Impressions are steady, clicks are down, and when you paste your target query into ChatGPT it does not cite you. But when you look at who it does cite, it is not a competitor brand. It is arxiv.org, or a Wikipedia article, or a Reddit thread from four years ago. This is a different problem from losing to another brand, and it calls for a different play.

Why AI defaults to arxiv, Wikipedia, and Reddit

The corpuses those sites represent were baked into the training data, and they continue to sit near the top of every engine’s trust hierarchy. An index of more than 680 million AI citations (5WPR, May 2026) found that fifteen domains account for 68% of all citations across ChatGPT, Claude, Perplexity, Gemini, and Google’s AI Overviews. Reddit alone accounted for about 40%, followed by Wikipedia, YouTube, LinkedIn, and Google’s own properties. Brand blogs are effectively absent from that top layer unless they earn a mention inside it.

When the citation on your target query is going to arxiv, that is a signal about the topic more than it is a signal about your content. The engine has decided that a research-grade source answers the question better than any brand asset would. Losing to a competitor brand means the citation window is open and you did not fill it. Losing to arxiv means the window is barely open at all.

Diagnose the pattern across four engines

Before choosing a play, run the same target query through all four engines and note what each cites. The pattern tells you what kind of loss you are looking at.

  • All four cite the same third party (arxiv, Wikipedia, Reddit). The answer sits fully outside the brand corpus for now. Winning the primary citation is unlikely; the secondary slot is the realistic target.
  • Two cite third parties, two cite competitor brands. Partial mix. Brand-earnable. Study what the two brand-cited pages do that yours does not.
  • One engine cites you, three do not. You are on the edge of the citation window. Usually a passage-structure or entity-signal fix, not a corpus problem.
  • None of the four cite anything on-topic. The query is new or under-answered. This is a rare and valuable position to fill.

Testing across ChatGPT, Claude, Google AI Overviews, and Perplexity in one sitting is essential. A win on one engine is not a win. AI engines diverge, and a play that lifts you in Perplexity can leave you invisible in ChatGPT.

Three plays when the winner is a third-party corpus

Play 1. Become the source the corpus cites. Publish research-grade content: original datasets, benchmarks, a whitepaper with a methods section, or a technical asset that a Wikipedia editor or a preprint author could reference. This is the slowest play and the most defensible. It works because the trust flows the right way: from your artifact into the corpus that AI already trusts.

Play 2. Meet the AI on the third-party corpus itself. Contribute expert-authored content to the platforms that already rank in the trust layer. That means Wikipedia edits with verifiable sources, quality Reddit answers written under a named account with real credentials, and LinkedIn posts that carry your entity signal. The loop AI closes here is community authority, not owned pages. When done consistently, your brand and your named experts start to appear inside the exact sources AI is already citing.

Play 3. Restructure your owned content for the secondary citation. Even when arxiv holds the primary slot, a well-structured brand page can win the second citation. Put the answer in the first sentence of each H2. Name the brand and category inside the answer paragraph, not only in the byline. Cite verifiable sources inline. Keep passages self-sufficient so they survive being scored in isolation. This is the fastest play, and it is entirely within your control.

What not to do

  • Do not chase backlinks in volume. AI engines score source quality, not link count. Ten mentions in low-trust directories do not move the citation.
  • Do not optimize for one engine at a time. Fixing a page until it wins in Perplexity while ignoring ChatGPT and Claude gets you a spot in one answer and a shrinking foothold everywhere else. The correct target is generalization across all four.
  • Do not confuse a citation with a recommendation. A citation quotes your page. A recommendation names your brand as the answer. Different mechanics, different work. The order matters: infrastructure first (citation), reputation second (recommendation).

How CLEO handles this

CLEO measures AI citation share across up to four engines active per site at a time at 95% Wilson Score confidence, so the cross-engine pattern above is visible on the dashboard rather than assembled by hand. When the diagnosis lands on Play 3 (passage restructuring), the fixes are deployed server-side directly into the CMS instead of routed through a development ticket queue.

CLEO’s own published case study for DisburseCloud documents a client’s AI citation share rising from 17% to 67% over 90 days at 95% Wilson confidence, with the loop between citation and content structured to compound. Read the full case study at regencleo.ai/case-studies/us-payment-engine.

About this article - Out-competed by arxiv.org in AI Answers: What Marketing Teams Do When the Citation Isn't a Competitor Brand

Your organic traffic is declining and the AI is citing arxiv.org, Wikipedia, or Reddit instead of your site. Why this pattern is different from losing to a competitor brand, how to diagnose it across ChatGPT, Claude, Google AI Overviews, and Perplexity, and the three plays marketing teams run in response.

Article details

Published July 30, 2026 by CLEO. Part of The Field Notes - the working journal of the CLEO Presence Engine at regencleo.ai/articles. Topics covered: GEO, AI Overviews, AI Citations, Citation Gap, Zero-Click Search.

Published on The Field Notes at regencleo.ai/articles. Learn more about the CLEO Presence Engine at regencleo.ai/engine. Methodology and scoring at regencleo.ai/methodology.