David Friedberg: AI Models are Training on YOUR DATA
David Friedberg raises concerns about LLM training on user conversation data without explicit consent or protection clauses. He provides anecdotal evidence that his organization's novel research insights appeared in subsequent model versions, suggesting their intellectual property was incorporated into model training without compensation or confidentiality agreements.
Summary
Friedberg discusses his concerns about trusting LLMs with proprietary knowledge, fearing that insights shared during conversations will be absorbed into model training. He describes a specific pattern: his team posed a novel scientific question to a model, received a response indicating the model hadn't considered this angle before. Later, using a different account with a newer model version, the same model described their exact previous discussion as if it were general training knowledge. While acknowledging these are anecdotal cases, Friedberg emphasizes his familiarity with his niche field, the novelty of their research, and the lack of published material containing these ideas—making it highly improbable that the model independently arrived at identical solutions through standard training data. He frames the core problem: models are legally permitted to train on identified data (data stripped of personal identifiers) but can still extract intellectual property value from conversations and repurpose those insights as generic training knowledge. Friedberg notes the absence of NDAs, confidentiality clauses, or IP protection agreements with service providers, which intensifies his concern about open-source models potentially gaining access to chat logs for training purposes. He views this as a mechanism through which competitive advantages derived from his organization's work could proliferate throughout the market without compensation or attribution.
About this episode
Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@allin Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg Intro Video Credit: https://x.com/TheZachEffect #allin #tech #news
Key Insights
- Friedberg observed that his organization's novel scientific insights appeared verbatim in subsequent model versions through different accounts, suggesting conversation data was directly incorporated into model training despite these insights having no published sources.
- Models can legally train on identified data (stripped of personal identifiers) while still extracting and commercializing intellectual property value from conversations, essentially converting users' proprietary insights into unmarked training material.
- Current service agreements with LLM providers lack NDAs, confidentiality clauses, or IP protection mechanisms, leaving organizations vulnerable to having their competitive insights distributed through open-source versions of models.
Topics
Transcript
[0:01] Should they trust any of these LLMs with their own knowledge, fearing it will be stolen into the core of the model? I've had experiences where we asked fairly new scientific questions and the model identified them as new insights. It's like, "Oh, I never thought about that ." "It would be interesting." Like, blah blah blah. And then, using a different account and requesting the same model later, or the next version, I encountered this. It's like, "Oh, well, you could do it like this." And it actually just [0:32] describes the exact thing we discussed in the previous version of the chat. Right now, these are just a few anecdotal cases. But I know the field we…
Full transcript available for MurmurCast members
Sign Up to AccessMore from All-In Podcast
Chamath: 3 Things Every Company Needs in the Age of Superintelligence
Chamath outlines three critical requirements for companies deploying superintelligence: end-to-end traceability to understand operations, mapping policies to risks, and maintaining audit evidence for accountability. He argues these practical governance measures are superior to delaying AI deployment through multinational regulation.
Jake Paul & The Chainsmokers: Turning Fame into Funds, Jake Enters Politics? & Venture Bubble Signs
A wide-ranging discussion featuring Jake Paul on entrepreneurship, boxing, and venture capital, followed by The Chainsmokers discussing their evolution from music producers to serious venture investors. Both segments explore how entertainment figures leverage fame and attention into sustainable business empires.
Daniel Ek: Life After Spotify, Broken Healthcare Incentives, Catching Disease Early & AI's Potential
Daniel Ek, Spotify's co-founder and now executive chairman, discusses his transition to healthcare through his new company Neko, which offers comprehensive preventive health screening for $499. The conversation covers Spotify's founding challenges, healthcare system incentive problems, and AI's potential in medical diagnostics and beyond.
Jason Calacanis: AI Is About to “Blow People’s Minds”
Jason Calacanis argues that AI products like Grok and Muse are reaching a tipping point where average consumers can derive real, tangible benefits for everyday tasks. He predicts that by year-end, ordinary people will adopt these tools as free personal assistants, marking a significant moment in AI democratization that could benefit lower-income Americans.
Jason Calacanis: Dario Seems to Be Sabotaging His Own IPO
Jason Calacanis critiques Anthropic CEO Dario Amodei for actions that appear to undermine the company's IPO prospects, specifically opening a lab after citing existential AI risks. Calacanis argues this represents a failure of leadership and suggests the board should replace the CEO with someone focused on running the company rather than sabotaging it.