Back to Blog
AI Search Strategy

Why AI Keeps Citing Reddit (and What Your Brand Should Do About It)

Reddit is one of the most-cited sources across ChatGPT, Perplexity and Google's AI. Here is why engines lean on it, the astroturfing trap, and the honest playbook.

The Diploria team

Ask an AI which product to buy in almost any category and there is a real chance the answer was quietly assembled from Reddit threads. Semrush data compiled for The Verge showed Reddit ranked as the most-cited domain across ChatGPT, Perplexity, Gemini, and Google AI Mode in May 2026, ahead of every news publisher and Wikipedia. Peec AI's 30-million-source analysis reached the same conclusion across five platforms. A pile of anonymous forum posts now outranks most of professional media as the raw material of AI answers.

That is strange enough to deserve a proper explanation, and consequential enough to deserve a plan. Here is both, including the part most coverage skips: the honest limits of the tactic.

How big is Reddit in AI search, really?

Big, but the size depends on how you count, and the spread is worth seeing before you reallocate budget.

At the high end, a widely circulated June 2025 study of 150,000+ citations found Reddit cited in 40.1% of AI answers across ChatGPT, Perplexity, Gemini, and Google AI Overviews. On Perplexity specifically, Reddit is the single largest source by a distance, at 46.7% of its top-10 citation share and by some counts about one in five of all its citations. On Google AI Overviews, trackers put it around 20%.

At the other end, Evertune's 200-million-prompt analysis, which measures share of total citations rather than share of answers, found no single domain, Reddit included, exceeds roughly 5% of all citations. And the number moves: Semrush's weekly tracking watched ChatGPT's Reddit citation rate collapse from about 60% of responses to about 10% in six weeks in late 2025, without an equivalent drop on other platforms.

So the accurate statement is not "Reddit is 40% of AI." It is: Reddit is persistently at or near the top of the source list on most engines, it is the backbone of Perplexity, and its weight can swing sharply per platform. That combination, high influence plus high volatility, is exactly why it needs monitoring rather than assumptions. The wider platform-by-platform picture is in our most-cited sources breakdown.

Why do AI engines lean on Reddit?

Four reasons, and none of them is an accident.

The engines are paying for it. Google signed a licensing deal with Reddit in early 2024, reported at roughly $60 million a year, giving it structured access to Reddit's content for training and surfacing. OpenAI signed its own partnership shortly after. When a platform pays for a corpus, that corpus tends to show up in the product.

It answers the questions people actually ask. AI queries are conversational and specific: "is X actually worth it," "X vs Y for a small team," "what do people regret about X." Almost nobody publishes professional content shaped like that. Reddit threads are shaped exactly like that, at enormous scale, with the long tail of niche questions no publisher would ever cover.

It reads as authentic experience. Engines increasingly weight first-hand experience over marketing language. A brand page says what the brand wants said. A thread with a hundred replies contains disagreement, caveats, and specifics, and the voting layer pre-ranks which answers a community found useful. For a model trying to synthesize "what do real users think," that structure is close to ideal.

It is fresh by default. Communities discuss products the week they change. Recency-weighted retrieval, which most engines now use, keeps pulling it in.

The catch: the astroturfing trap

Everything above is also visible to every growth hacker on earth, which is why the space between "participate on Reddit" and "manipulate Reddit" has become the most debated question in AI visibility.

Be clear-eyed about the risks. Reddit's own enforcement is aggressive: the platform removes on the order of 25,000 spam posts a day and blocks millions of spam views, and communities ban brands faster than admins do. Fake accounts praising your product are detectable, deletable, and reputationally radioactive if exposed, and a thread about your astroturfing can itself become the content AI cites about you. There is also a quality problem worth respecting: upvotes measure popularity, not accuracy, and an engine citing a thread implies relevance, not verification. Building your brand's AI presence on manipulated threads is building on sand that is also flammable.

What brands should actually do

The legitimate playbook is unglamorous, and it works better than the shortcuts.

Find out where Reddit already talks about you. Before creating anything, learn which threads and subreddits engines are pulling from when they answer questions in your category, and what those threads say. This is measurement, not marketing, and it regularly surfaces the exact thread feeding a wrong claim or a competitor's win.

Show up transparently where you have standing. Answer questions about your own product with disclosure, fix misinformation politely with sources, and contribute expertise in your domain beyond your own brand. Communities tolerate and often welcome identified brand representatives who are useful. They exile shills.

Give communities something worth discussing. The durable way to earn Reddit presence is the same as earning press: ship things, publish original data, and be genuinely notable. Threads about you that you never touched are the strongest signal there is.

Do not buy posts, seed fake reviews, or run sockpuppets. Beyond the ethics, it is strategically self-defeating: you would be optimizing a trust signal by destroying the trust.

Watch it like a channel, because it is one. Reddit's citation weight moves per engine per quarter. Treat the threads that feed AI answers about your category as inventory to monitor, the same way you monitor rankings.

This is precisely the workflow Diploria's Question Radar was built for: it discovers the questions being asked about your brand and category on Reddit, alongside AI fan-out queries and keyword sources, so you can track the ones you are currently blind to. Paired with the Citations view, you can see which specific threads engines cite in answers about you, and whether that changes after you engage. Start with the free AI visibility check to see whether Reddit is already shaping your answers.

FAQs

Frequently asked questions

A combination of licensing (OpenAI has a data partnership with Reddit), retrieval preferences that reward first-hand experience and freshness, and fit: conversational AI queries look like forum questions, and Reddit is the largest corpus of answered forum questions on the open web.

Know exactly where
AI mentions your brand

Track ChatGPT, Gemini, Perplexity, Claude and 11 more AI engines - weekly. See how your competitors rank, spot gaps, and fix them fast.