Ask an AI assistant almost anything practical and there is a decent chance a Reddit thread shaped the answer. That is not an accident. Reddit signed licensing deals with Google and OpenAI worth hundreds of millions, and independent studies keep finding it among the most-cited domains in AI answers. Forums won because they hold what machines cannot fake: first-hand human experience, arranged as questions and answers. The move for you is participation, not promotion.
For twenty years, the joke was that the best search engine was Google with "reddit" typed after your query. The machines were listening. Now they do it for you, invisibly, inside almost every answer.
Why does ChatGPT keep quoting a forum?
Because the forum holds the thing every answer engine is desperate for: honest, specific, first-hand experience, written by people with no commercial reason to lie. When someone asks an assistant "is this course worth it" or "which accountant handles contractors well", the most useful raw material on earth is a thread where thirty real people compared notes. Marketing copy tells the machine what a company wishes were true. A forum tells it what customers found out.
The scale of this preference is measurable. A Semrush study of millions of citations found Reddit among the most-cited domains across AI platforms, and separate research covered by Search Engine Land found AI search engines cite Reddit, YouTube, and LinkedIn more than almost anything else (full study link in the sources below). The exact share swings month to month and platform to platform, but the direction has been unmistakable: community content is not a fringe source. It is a load-bearing wall of the answer layer.
What did the licensing deals actually buy?
Money formalised the relationship. In February 2024, on the same day it filed for its IPO, Reddit announced a content licensing agreement with Google, reported at roughly $60 million a year, giving Google structured, real-time access to its forums. Three months later came a partnership with OpenAI, announced in May 2024, granting access to what the companies called "real-time, structured and unique content" from Reddit. The IPO prospectus put a number on the whole programme: data licensing contracts with an aggregate value of $203 million, on terms of two to three years.
| When | What happened | What it means |
|---|---|---|
| Feb 2024 | Reddit announces a licensing deal with Google, reported at about $60M a year | Google gets structured, real-time forum content for training and products |
| Mar 2024 | IPO prospectus discloses $203M in aggregate data licensing contracts, 2-3 year terms | Community conversation is formally priced as an AI asset |
| May 2024 | Reddit announces a partnership with OpenAI | ChatGPT gains access to real-time Reddit content, and Reddit gains AI features |
Sources: Columbia Journalism Review; TechCrunch, May 2024.
Notice what was actually sold. Not ad space, not user data in the surveillance sense, but conversation itself: the accumulated question-and-answer record of twenty years of humans helping each other. When that record becomes one of the most expensive datasets in the AI economy, it tells you precisely what the machines consider scarce. They can generate prose all day. They cannot generate having-been-there.
There is a lesson in the pricing, too. The deals value conversation as an ongoing feed, not a one-off archive: real-time access, renewable terms, and reported discussions about pricing that rises as the data proves more useful inside answers. The engines are not buying Reddit's past. They are subscribing to its present, which means the threads being written this month are already tomorrow's answer material.
What makes forum answers so liftable?
Three structural properties, each worth understanding because each one is imitable in your own publishing.
- Threads are question-shaped. A Reddit thread begins with exactly the kind of messy, situational question users type into assistants, and the replies are candidate answers. For a machine assembling a response, this is pre-chewed food. Ordinary web pages have to be reverse-engineered into answers; threads already are answers.
- Votes are a quality signal. Upvotes, downvotes, and moderation give engines a crude but useful proxy for which answer a community found credible. It is noisy and gameable, but it is more signal than a random page provides, and machines take whatever signal they can get.
- Experience is stated in the first person. "I ran this for six months and here is what broke" is the experience component of E-E-A-T made flesh. It is specific, dated, checkable against other replies, and impossible to fake at scale without a community noticing.
This is the same logic behind the publishing AI actually reads: machines lift what is specific, structured, and corroborated. Forums simply industrialised those properties before anyone knew machines would be the audience.
Should you go and post on Reddit, then?
Carefully, and only in the spirit the place demands. The worst possible reading of this essay is "great, I shall go and spam twelve subreddits with my name". Communities have spent two decades sharpening their immune systems against exactly that. Self-promotion gets deleted, downvoted, and occasionally screenshotted for public ridicule, and a deleted post is invisible to machines and humans alike.
The move that works is slower and more honest. Pick the two or three communities where your actual expertise is useful. Answer real questions with real substance, under a consistent username that connects, somewhere, to your real identity. Cite your sources, admit uncertainty, and let the votes do their work. Over months, two things happen: humans in your field start recognising the name, and the threads you contributed to become part of the corpus the engines draw on. You are not gaming the system. You are becoming part of the record the system trusts.
Machines quote communities because communities cannot be bought, only earned. The moment your participation looks like marketing, the community rejects it, and the machine never sees it. Contribution is the only currency that clears.
Which communities count beyond Reddit?
Reddit is the headline because it signed the headline deals, but the machine's appetite is for candid conversation wherever it accumulates in crawlable text. Stack Exchange sites carry enormous weight in technical answers for the same structural reasons: question-shaped threads, voting, and an obsessive culture of citation. Quora persists in AI answers well past its cultural peak because its archive is vast and question-formatted. Hacker News shapes answers about startups and engineering. Industry-specific forums, the quiet ones where actuaries or dermatologists or timber-frame builders compare notes, often dominate answers in their niches precisely because nothing else on the open web discusses those questions honestly.
Two absences are worth noting. Discord and Slack communities, where much of the sharpest professional conversation moved, are largely invisible to engines: private, unindexed, ephemeral. Wisdom that lives there does not feed the answer layer at all. And LinkedIn sits in a middle state, partially crawlable and heavily cited in professional queries, which is one reason a consistent presence there still earns its keep. The strategic reading is simple: of all the places your expertise currently lives, only the indexable ones are building your machine reputation. If everything you know is spent in private channels, the engines will go on believing you know nothing.
So audit where your field's honest conversation is crawlable. One good answer on a public forum, findable and dated and signed, outworks fifty brilliant messages in a channel no machine will ever read.
Does this replace your own website?
No, and the relationship is worth getting right. Community mentions are corroboration: third parties discussing you in a venue you do not control, which is precisely why engines weight them. But corroboration needs something to corroborate. If a thread mentions your name and the machine goes looking for you, it needs to find a coherent entity: a site that states who you are, work published under your byline, profiles that agree with each other. We have written about where AI looks before recommending someone, and the honest picture is a mesh: your own property at the centre, licensed community platforms and reference sites like Wikipedia and Wikidata supplying the trust around it.
One caution belongs here. Citation patterns churn: roughly 40-60% of sources cited in AI answers rotate month over month, according to Semrush's AI Visibility Index. Reddit's share in specific engines has swung dramatically within single quarters. So do not build your whole visibility on any single platform's current weighting, forums included. Build the durable thing, a verifiable identity plus genuine community standing, and let the weightings swing around you.
A worked example: one honest answer, months later
Here's a simple, hypothetical illustration of how this plays out. Someone genuinely knowledgeable about small business accounting spends fifteen minutes one evening answering a stranger's question in a small business subreddit, thoroughly, with specifics, under a username loosely tied to their actual site. The post gets a modest number of upvotes, nothing viral, and they forget about it within a week. Months later, someone else asks an AI assistant a very similar question, and the assistant's answer echoes the structure and some of the specific advice from that old thread, occasionally crediting the general consensus of "several small business owners" without naming anyone directly. The original poster may never know their answer influenced anything. But if their identity was even loosely connected, consistent username, a linked profile, the same person answering similar questions across a few threads over time, the credit compounds slowly into exactly the kind of recognized presence this whole site is about. One answer rarely does it. A pattern of honest answers, over time, does.
What does the front page of the machine look like?
Something rather beautiful, if you squint. The old front page of the internet was an editorial artefact: a home page, a headline stack, someone deciding what mattered. The new one is a synthesis of conversations, assembled per question, per person, in real time. The sources that feed it are wherever humans talk candidly at scale, which today means Reddit and its cousins, plus the podcast interviews that keep surfacing in answers, plus whatever community your niche actually lives in.
For anyone building a name, the instruction hiding in all this is almost old-fashioned: be genuinely useful where your people gather, and keep your identity consistent enough that the credit accrues to you. The engines did not invent that advice. They just started enforcing it. If you want to know which communities matter for your field and how your name currently reads to the machines, that audit is the first thing we run in our engagements.
Questions people ask
Why does AI cite Reddit so much? +
Did Reddit really sell its data to AI companies? +
Should I promote myself on Reddit? +
Do private Discord or Slack groups help my AI visibility? +
Does one good forum answer make a real difference? +
Should I stop investing in my own website if forums matter this much? +
Sources
- TechCrunch: OpenAI inks deal to train AI on Reddit data (May 2024)
- Columbia Journalism Review: Reddit is winning the AI game (licensing deal values and IPO disclosures)
- Semrush: The most-cited domains in AI, a 3-month study
- Search Engine Land: AI search engines cite Reddit, YouTube, and LinkedIn most
- Similarweb: Gen AI stats (citation rotation, Semrush AI Visibility Index)
Curious what AI says about you?
Start with a check-up. We'll show you the exact words the engines return about your name, then map the fastest signal to move.
Say my name →