Episode 168 debate report.

Share

Featuring

Chamath Palihapitiya Jason Calacanis David Sacks David Friedberg
Episode 168 video thumbnail

Episode 168 ricochets from Gemini's launch disaster and the case against bloated HR teams to Klarna's AI support bot and Reddit's IPO math. The real spice arrives when Chamath declares a new "TAC 2.0" royalty stream for the web and Friedberg asks whether most training data will age more like yesterday's box score than *The Sopranos*. Chamath and Friedberg split the win: one nailed the recurring-access mechanism, while the other nailed how selective and fragile the market would remain.

Spice rack

🌶️ 🌶️ Medium heat 00:43:08

Will AI data licensing become recurring 'TAC 2.0' revenue, or remain a selective and uncertain content market?

Original point: Google was beginning to pay platforms for training data just as it paid distribution partners for search traffic, creating a new high-margin revenue stream he called 'TAC 2.0.'

What everyone argued

Chamath Palihapitiya

Chamath argued that owners of unique, continuously refreshed datasets could license them to model builders as an incremental, high-margin revenue stream. He expected the model to spread from giants such as Reddit to smaller sites once buyers learned how to attribute each dataset's incremental value, while acknowledging that small publishers might earn only modest amounts.

Jason Calacanis

Jason leaned toward durable licensing value for platforms with ongoing content, while distinguishing valuable archives from disposable old material. He argued that selected bodies of historical interviews or media could remain useful even as routine web content decayed, and floated acquisition as another path for strategically important data platforms.

David Sacks

Sacks argued that content owners faced an unfavorable market unless they coordinated: many suppliers were competing for a small number of model buyers, which would suppress prices. He compared the missing coordination to music licensing and said a federation could strengthen publishers' bargaining power.

David Friedberg

Friedberg rejected the simple TAC analogy. He argued that model builders were buying selected content more like streaming services license programming: buyers would learn which datasets add value, old material would often become stale, and contracts could be continuous or chunky depending on the corpus. He therefore treated the long-run monetization model as unsettled.

Winner circle

Chamath Palihapitiya David Friedberg

Chamath was right that fresh, differentiated data could become recurring access revenue, and Reddit's filings validate that mechanism unusually clearly. Friedberg was right that 'TAC 2.0' oversold the breadth and predictability of the market: licensing remained negotiated, concentrated, and dependent on continuing incremental value. This is a co-win because Chamath got the revenue architecture right while Friedberg got the market's selectivity and durability risk right; neither broadest version survives hindsight.

Commentary

Chamath Palihapitiya

Commentary

Chamath identified the recurring-access mechanism correctly before the market had much history. His weak point was breadth: evidence for a few differentiated, frequently updated corpora does not establish a web-wide TAC system.

Assumptions and fact checks
Assumptions
Agree
Assumption

Unique data will remain valuable enough after initial model training to support recurring payments.

Why it matters

Fresh access mattered in the signed Reddit contracts, and both Google and OpenAI emphasized current, structured content. Continued value depends on freshness and product use, not merely ownership of an old archive.

Neutral
Assumption

A meaningful automated licensing market will extend to small websites and apps.

Why it matters

The mechanism is plausible, but Reddit says substantially all licensing contract value still comes from two partners. There is not enough evidence that small publishers gained a material, standardized revenue channel.

Fact checks
True High confidence
Claim

Reddit disclosed data-licensing arrangements worth $203 million with terms of two to three years.

Check

Reddit's IPO filing states that January 2024 data-licensing arrangements had an aggregate contract value of $203.0 million, ran for two to three years, and were expected to produce at least $66.4 million of 2024 revenue.

Sources [1]
True High confidence
Claim

The Reddit arrangements involved continuing delivery of data rather than only a one-time archive transfer.

Check

The IPO filing describes continuous Data API access plus quarterly transfers over the contract term. Reddit's Google announcement likewise describes programmatic access to constantly evolving posts and comments.

Sources [1] [2]

Jason Calacanis

Commentary

Jason's archive-versus-feed distinction sharpened the debate, and his concession showed good discipline. The takeover talk was the weakest part because it skipped the price, exclusivity, antitrust, and integration tradeoffs facing buyers.

Assumptions and fact checks
Assumptions
Agree
Assumption

Frequently refreshed platforms and a limited set of culturally important archives retain licensing value longer than generic old web content.

Why it matters

Google and OpenAI explicitly valued Reddit's current, structured conversations, while other multi-year publisher agreements have covered both current and archived material. The distinction is more useful than treating all data as one commodity.

Disagree
Assumption

Data-rich platforms such as Reddit, Quora, and Stack Overflow would be acquired because their corpora became strategically indispensable.

Why it matters

The prediction was too confident and ignored the cheaper option of nonexclusive term licenses. Reddit remained independent through the hindsight window, and Stack Overflow had already been acquired by Prosus in 2021 for reasons predating this AI-licensing wave.

David Sacks

Commentary

Sacks contributed the debate's cleanest account of bargaining power. He would have strengthened it by separating collective licensing from a coordinated threat to de-index, which carries different legal, traffic, and business risks.

Assumptions and fact checks
Assumptions
Agree
Assumption

Without collective bargaining or uniquely scarce data, many content suppliers and few model buyers will push licensing prices down.

Why it matters

Reddit's dependence on two partners and its warning that some model builders may use open content without paying support the bargaining-power concern. Truly differentiated real-time datasets can still escape commodity pricing.

Neutral
Assumption

A publisher federation could materially improve licensing leverage.

Why it matters

Pooling rights can improve leverage, but publishers differ in rights ownership, content quality, update frequency, and willingness to withhold indexing. The episode did not show how a federation would solve those coordination and enforcement problems.

David Friedberg

Commentary

Friedberg supplied the most important limit on the slogan: ongoing payment follows ongoing differentiated value, not mere content ownership. His argument would have been cleaner without unsupported, contradictory estimates of how much data humanity creates each day.

Assumptions and fact checks
Assumptions
Agree
Assumption

Most older content loses incremental model value as new data accumulates and buyers learn which sources matter.

Why it matters

Generic or duplicated data is substitutable, while filings and partnerships emphasize differentiated and continually updated access. Some high-quality archives remain exceptions, as Jason noted.

Agree
Assumption

The market would remain selective rather than becoming an automated TAC-like payment layer for the whole web.

Why it matters

By year-end 2025 Reddit still described content licensing as early, concentrated in two partners, and vulnerable to nonrenewal. That is a negotiated market for scarce corpora, not a broad attribution system.

Fact checks
Unclear High confidence
Claim

Google's network advertising business was making about $10 billion per quarter at the time.

Check

Alphabet reported $31.3 billion of Google Network revenue for all of 2023, or about $7.8 billion per quarter on a simple average. The broader point that most network-ad revenue is paid to partners is supported by Alphabet, but the revenue figure was overstated.

Sources [1]