Episode 168 ricochets from Gemini's launch disaster and the case against bloated HR teams to Klarna's AI support bot and Reddit's IPO math. The real spice arrives when Chamath declares a new "TAC 2.0" royalty stream for the web and Friedberg asks whether most training data will age more like yesterday's box score than *The Sopranos*. Chamath and Friedberg split the win: one nailed the recurring-access mechanism, while the other nailed how selective and fragile the market would remain.
Spice rack
Will AI data licensing become recurring 'TAC 2.0' revenue, or remain a selective and uncertain content market?
Original point: Google was beginning to pay platforms for training data just as it paid distribution partners for search traffic, creating a new high-margin revenue stream he called 'TAC 2.0.'
What everyone argued
Chamath Palihapitiya
Chamath argued that owners of unique, continuously refreshed datasets could license them to model builders as an incremental, high-margin revenue stream. He expected the model to spread from giants such as Reddit to smaller sites once buyers learned how to attribute each dataset's incremental value, while acknowledging that small publishers might earn only modest amounts.
Jason Calacanis
Jason leaned toward durable licensing value for platforms with ongoing content, while distinguishing valuable archives from disposable old material. He argued that selected bodies of historical interviews or media could remain useful even as routine web content decayed, and floated acquisition as another path for strategically important data platforms.
David Sacks
Sacks argued that content owners faced an unfavorable market unless they coordinated: many suppliers were competing for a small number of model buyers, which would suppress prices. He compared the missing coordination to music licensing and said a federation could strengthen publishers' bargaining power.
David Friedberg
Friedberg rejected the simple TAC analogy. He argued that model builders were buying selected content more like streaming services license programming: buyers would learn which datasets add value, old material would often become stale, and contracts could be continuous or chunky depending on the corpus. He therefore treated the long-run monetization model as unsettled.
Winner circle
Chamath was right that fresh, differentiated data could become recurring access revenue, and Reddit's filings validate that mechanism unusually clearly. Friedberg was right that 'TAC 2.0' oversold the breadth and predictability of the market: licensing remained negotiated, concentrated, and dependent on continuing incremental value. This is a co-win because Chamath got the revenue architecture right while Friedberg got the market's selectivity and durability risk right; neither broadest version survives hindsight.
Commentary
Chamath Palihapitiya
Assumptions and fact checks
Unique data will remain valuable enough after initial model training to support recurring payments.
Why it mattersFresh access mattered in the signed Reddit contracts, and both Google and OpenAI emphasized current, structured content. Continued value depends on freshness and product use, not merely ownership of an old archive.
A meaningful automated licensing market will extend to small websites and apps.
Why it mattersThe mechanism is plausible, but Reddit says substantially all licensing contract value still comes from two partners. There is not enough evidence that small publishers gained a material, standardized revenue channel.
Reddit disclosed data-licensing arrangements worth $203 million with terms of two to three years.
CheckReddit's IPO filing states that January 2024 data-licensing arrangements had an aggregate contract value of $203.0 million, ran for two to three years, and were expected to produce at least $66.4 million of 2024 revenue.
The Reddit arrangements involved continuing delivery of data rather than only a one-time archive transfer.
CheckThe IPO filing describes continuous Data API access plus quarterly transfers over the contract term. Reddit's Google announcement likewise describes programmatic access to constantly evolving posts and comments.
Jason Calacanis
Jason's archive-versus-feed distinction sharpened the debate, and his concession showed good discipline. The takeover talk was the weakest part because it skipped the price, exclusivity, antitrust, and integration tradeoffs facing buyers.
Assumptions and fact checks
Frequently refreshed platforms and a limited set of culturally important archives retain licensing value longer than generic old web content.
Why it mattersGoogle and OpenAI explicitly valued Reddit's current, structured conversations, while other multi-year publisher agreements have covered both current and archived material. The distinction is more useful than treating all data as one commodity.
Data-rich platforms such as Reddit, Quora, and Stack Overflow would be acquired because their corpora became strategically indispensable.
Why it mattersThe prediction was too confident and ignored the cheaper option of nonexclusive term licenses. Reddit remained independent through the hindsight window, and Stack Overflow had already been acquired by Prosus in 2021 for reasons predating this AI-licensing wave.
David Sacks
Sacks contributed the debate's cleanest account of bargaining power. He would have strengthened it by separating collective licensing from a coordinated threat to de-index, which carries different legal, traffic, and business risks.
Assumptions and fact checks
Without collective bargaining or uniquely scarce data, many content suppliers and few model buyers will push licensing prices down.
Why it mattersReddit's dependence on two partners and its warning that some model builders may use open content without paying support the bargaining-power concern. Truly differentiated real-time datasets can still escape commodity pricing.
A publisher federation could materially improve licensing leverage.
Why it mattersPooling rights can improve leverage, but publishers differ in rights ownership, content quality, update frequency, and willingness to withhold indexing. The episode did not show how a federation would solve those coordination and enforcement problems.
David Friedberg
Friedberg supplied the most important limit on the slogan: ongoing payment follows ongoing differentiated value, not mere content ownership. His argument would have been cleaner without unsupported, contradictory estimates of how much data humanity creates each day.
Assumptions and fact checks
Most older content loses incremental model value as new data accumulates and buyers learn which sources matter.
Why it mattersGeneric or duplicated data is substitutable, while filings and partnerships emphasize differentiated and continually updated access. Some high-quality archives remain exceptions, as Jason noted.
The market would remain selective rather than becoming an automated TAC-like payment layer for the whole web.
Why it mattersBy year-end 2025 Reddit still described content licensing as early, concentrated in two partners, and vulnerable to nonrenewal. That is a negotiated market for scarce corpora, not a broad attribution system.
Google's network advertising business was making about $10 billion per quarter at the time.
CheckAlphabet reported $31.3 billion of Google Network revenue for all of 2023, or about $7.8 billion per quarter on a simple average. The broader point that most network-ad revenue is paid to partners is supported by Alphabet, but the revenue figure was overstated.

Chamath identified the recurring-access mechanism correctly before the market had much history. His weak point was breadth: evidence for a few differentiated, frequently updated corpora does not establish a web-wide TAC system.