Reddit sues Anthropic as Huffman claims a third of OpenAI's training set
Article excerpt
Reddit chief executive Steve Huffman used a June 22, 2026 podcast recording, released July 3, to put a number on how much of modern language models rests on his platform, and to explain why the company that once gave its data away now litigates over it. The interview appeared on Mixed Signals, the media podcast produced by Semafor, and was published on YouTube on July 3, 2026. Media editor Max Tani and editor-in-chief Ben Smith recorded it at Cannes Lions, on the day Reddit turned 21. Huffman co-founded the site in 2005 and launched it on June 22 of that year, at the age of 21. "So today is Reddit's 21st birthday," he said, noting that the platform now accounts for exactly half his life. Much of the conversation concerned a question that has become commercially significant for anyone buying media against user-generated content: how much of what large language models produce originates on Reddit, and what that content is worth. Asked how much of the output from Gemini, Claude or ChatGPT traces back to Reddit, Huffman was blunt. "A lot. We don't know for sure and they don't tell us, but it's a lot," he said. He then supplied the only public anchor he has. "One data point I have is OpenAI's last public research paper," he said, adding that it "said that Reddit was about a third of their training set." He placed that paper at the GPT-2 or GPT-3 stage and acknowledged that...
Keep reading with a free account
The rest of this article, and every signal for Reddit, is in your free account.
Extracted from this sentence
Reddit reported 726 million dollars in fourth-quarter 2025 revenue with advertising up 75% , then 625 million dollars in first-quarter 2026 advertising revenue, a 74% year-on-year increase, alongside 127 million daily active uniques.
