by datastudy.nl

Tuesday, August 4, 2026

Business

Reddit AI search spam hits 25,000 poisoned posts daily

Reddit AI search spam floods the platform as brands chase chatbot citations. Reddit catches 25,000 spammy posts daily and blocks 23 million views.

Reddit AI search spam enforcement scale: 23 million spam views blocked daily narrowing to 25,000 spammy posts and comments caught daily, showing the gap between visibility blocked and content removed.
Reddit blocks 23 million spam views and catches 25,000 spammy posts and comments per day. Source: Reddit via The Verge.

Reddit AI search spam has reached industrial scale. The platform that users once fled to for authentic discussion is now the most-cited domain by AI chatbots, surpassing every news publisher, scholarly repository, and even Wikipedia in May 2026, according to data compiled by Semrush for The Verge. That distinction has turned Reddit into a battleground. Marketers, brands, and stealth agencies are flooding subreddits with coordinated promotional content designed to manipulate what ChatGPT, Perplexity, and Google's Gemini surface in their answers. Reddit says it is now catching 25,000 spammy posts and comments and blocking 23 million spam views every day. The unpaid volunteer moderators who run individual communities are on the front line of a fight they never signed up for: defending the authenticity that makes Reddit valuable to AI in the first place.

What is happening to Reddit right now?

The Semrush data, compiled for The Verge, found Reddit was the most-cited domain in May 2026 across ChatGPT, Perplexity, and Google's Gemini and AI Mode. More than any news publisher, more than Wikipedia, more than academic repositories. This happened because users spent years appending "reddit" to search queries to escape SEO junk, and Google responded by ranking Reddit prominently in results.

Now that AI search tools cite Reddit content in their answers, brands have realized they do not even need a backlink anymore. A mere mention in a Reddit thread is enough to get a brand surfaced by a chatbot. The old SEO playbook required getting links pointing to your site. The new playbook, sometimes called "answer engine optimization" or "generative engine optimization," just requires getting your brand name into a conversation that an LLM might scrape and cite.

The result is coordinated spam campaigns across subreddits, with accounts posting seemingly authentic questions and answers that are actually designed to inject brand mentions into AI training data and retrieval pipelines.

Reddit itself acknowledged the problem in early July 2026, announcing it was beefing up spam detection tools with AI and LLMs to catch what the company called "highly subtle, coordinated patterns of fake behavior and artificial hype." Jared Nelson, Reddit's senior director of global corporate reputation communications, told The Verge that "every time bad actors change their playbook, we strengthen ours." Nelson declined to name specific detection signals, saying Reddit did not want to give bad actors a playbook for evading detection.

How are brands gaming the system?

The tactics have evolved beyond obvious spam. A suspected marketer might pose as a genuine user asking "How does everyone keep track of their skincare routine?" and then a wave of comments mention specific AI skincare apps. Other posts follow a self-deprecating format: "I thought this product would be too aggressive, but it was actually very gentle." The posts mimic how a user might prompt a chatbot, because that is exactly what they are designed to influence.

The Verge's reporting details a case from r/SkincareAddiction, a subreddit with more than 1 million weekly visitors. A user named Primary-Taro4254 recommended a hypochlorous acid spray from Honeydew Labs across multiple unrelated threads and subreddits, complete with specific concentration details and credential name-drops. A community member caught on, alerted moderators, and the account was filtered. This is one account among thousands.

Moderators across the platform report similar patterns. Users on r/CleaningTips debated banning product names entirely. r/BuyItForLife was warned it was being "brigaded by ad bots to skew Google and AI search results." r/FoodNYC specifically calls out brands targeting and commenting on old posts. A moderator of r/salestechniques posted a "name and shame" thread of brands astroturfing the subreddit. A separate TechTimes report on the same phenomenon characterized these as "poisoned posts" designed to target ChatGPT specifically.

Maya Adivi, a moderator of r/SkincareAddiction, told The Verge that "with AEO right now, no one knows exactly how to do it, and everyone tells you how important Reddit is." She described the spam as "a lot more subtle" than what moderators faced years ago, when telltale signs included tracking tags in product links or before-and-after photos lifted directly from product listings.

The answer comes down to what AI tools use as ground truth. When ChatGPT or Perplexity generates an answer about which product to buy or which service to trust, it pulls from sources it deems authoritative and authentic. Reddit, with its pseudo-anonymous users and community-moderated discussions, has become the closest thing to a crowdsourced trust signal that AI can scrape at scale.

Reddit says roughly 40 percent of discussions on the platform are commercial in nature, according to Nelson, as the chart below breaks down. That is a massive surface area for brands to target.

Donut chart of Reddit discussion types: commercial at 40 percent, non-commercial at 60 percent.
Reddit says roughly 40 percent of platform discussions are commercial in nature; the remaining 60 percent are non-commercial. Source: Reddit spokesperson via The Verge.

Google spokesperson Jennifer Kutz told The Verge that Reddit "gets no special preference" in Google's ranking system and that Google has kept results 99% spam-free for years. But the Semrush data tells a different story about what actually shows up in AI answers. Semrush also found that compared to a year ago, Google is inserting AI Overviews much more frequently into search queries with commercial intent, like when shoppers research products or compare brands. The pipeline from Reddit thread to AI answer to consumer decision is getting shorter and wider.

For anyone building products that rely on AI search visibility, this is the new terrain. The platforms that AI cites most are the ones most targeted by manipulation, and Reddit is at the top of that list. As Cloudflare's changes to agent crawler permissions show, the rules of access to web content for AI are in flux, and Reddit's spam fight is part of the same broader collision between AI retrieval and content authenticity.

What does this mean for builders and marketers?

For a builder or founder trying to get visibility in AI search, the implications are concrete:

  • Mentions matter more than links. If your product gets named in a Reddit thread that an AI tool scrapes, that mention can surface in chatbot answers even without a link to your site. The backlink economy is being supplemented by a mention economy, and the mention economy is easier to game.
  • Authenticity is the ranking signal, and it is under active attack. The value of Reddit to AI comes from perceived authenticity. As spam erodes that, the signal degrades. If you are building retrieval systems or training on Reddit data, you need to account for the contamination.
  • Moderators are the last line of defense, and they are unpaid. The communities that AI tools cite are curated by volunteers. If you are building AI products that depend on user-generated content, you are depending on moderation infrastructure you do not control and do not pay for.
  • 40 percent commercial discussion means 40 percent attack surface. Anywhere users talk about products, marketers will try to insert themselves. If your product surfaces recommendations from Reddit or similar forums, you need spam detection in your pipeline.

For teams building RAG systems or fine-tuning on web data, the Reddit spam problem is a data quality problem. If your retrieval pipeline pulls from Reddit, you are pulling from a source under active coordinated manipulation. The 25,000 posts and comments Reddit catches daily are the ones it knows about. The ones it misses are in your training data and your retrieval results.

Mike Moschella, director of analytics at marketing and PR firm DKC, told The Verge that "you can't bullshit Reddit." He argues brands should make their pitch more clearly and specifically rather than repeating lofty promises. The communities that catch spam are sophisticated, and getting caught means becoming "persona non grata," as Adivi puts it. She was recently alerted to a marketing agency that claimed it had "hijacked" r/SkincareAddiction to promote a Korean beauty brand. She removed the thread and added the brand to the keyword filtering list.

Should you invest in Reddit as a channel?

This depends on what you are building. If your product is consumer-facing and people ask AI tools about it, Reddit visibility matters. But the tactics matter enormously.

Some brands have taken to creating their own subreddits where they moderate the community directly. This is legitimate but limited: a brand subreddit has less authority than an independent community discussing your product organically. Reddit only introduced native tools to help businesses grow their presence on the platform in 2024, making this a very young channel.

The honest read for builders is this: Reddit is a high-signal environment for AI search, and that signal is being actively polluted. If your roadmap depends on AI search visibility, you need to think about content authenticity the way you think about code quality. Sloppy astroturfing will get you banned and potentially named and shamed in a moderator thread. Genuine participation, with disclosure, is the only sustainable approach.

For teams building on Reddit's data, a Mashable report notes that the spam tactics are specifically designed to be subtle enough to evade both human moderators and automated detection. Your data pipeline needs the same skepticism that the best moderators apply manually.

What comes next for AI search integrity?

Reddit announced in early July 2026 that it is using AI and LLMs to catch coordinated manipulation. The company evaluates accounts at creation, looks for coordinated behavior across accounts and conversations, and uses platform activity and user reports to improve detection. Reddit says it is catching 25,000 spammy posts and comments and blocking 23 million spam views a day as a result.

The arms race mirrors the broader fight over AI training data quality. As research on chain-of-thought forgery bypassing LLM guardrails demonstrates, even safety mechanisms built into language models can be circumvented with the right approach. Reddit's spam problem is a preview of what every platform that contributes to AI training and retrieval will face.

The deeper question is whether the authenticity signal that makes Reddit valuable to AI can survive the attention it attracts. Google says Reddit gets no special preference. The Semrush data says Reddit is the most cited domain. Those two statements can both be true and still point to a system under strain. If AI search tools continue to cite Reddit disproportionately, the incentive to game Reddit will only grow, and the cost of defending it will fall on the same unpaid moderators who have been doing the work for years.

The authenticity paradox

The irony is brutal. Reddit became the most trusted source in AI search because it felt unmanipulated. Every brand that successfully games a subreddit to get cited by a chatbot makes the next citation worth less. The platforms, the brands, the moderators, and the AI tools are all betting on the same scarce resource: authenticity. Whoever depletes it fastest loses the most.

Sources