OpinionDiscussion

David Friedberg: AI Models are Training on YOUR DATA

All-In Podcast1m 42s

David Friedberg raises concerns about LLM training on user conversation data without explicit consent or protection clauses. He provides anecdotal evidence that his organization's novel research insights appeared in subsequent model versions, suggesting their intellectual property was incorporated into model training without compensation or confidentiality agreements.

Summary

Friedberg discusses his concerns about trusting LLMs with proprietary knowledge, fearing that insights shared during conversations will be absorbed into model training. He describes a specific pattern: his team posed a novel scientific question to a model, received a response indicating the model hadn't considered this angle before. Later, using a different account with a newer model version, the same model described their exact previous discussion as if it were general training knowledge. While acknowledging these are anecdotal cases, Friedberg emphasizes his familiarity with his niche field, the novelty of their research, and the lack of published material containing these ideas—making it highly improbable that the model independently arrived at identical solutions through standard training data. He frames the core problem: models are legally permitted to train on identified data (data stripped of personal identifiers) but can still extract intellectual property value from conversations and repurpose those insights as generic training knowledge. Friedberg notes the absence of NDAs, confidentiality clauses, or IP protection agreements with service providers, which intensifies his concern about open-source models potentially gaining access to chat logs for training purposes. He views this as a mechanism through which competitive advantages derived from his organization's work could proliferate throughout the market without compensation or attribution.

About this episode

Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@allin Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg Intro Video Credit: https://x.com/TheZachEffect #allin #tech #news

Key Insights

  • Friedberg observed that his organization's novel scientific insights appeared verbatim in subsequent model versions through different accounts, suggesting conversation data was directly incorporated into model training despite these insights having no published sources.
  • Models can legally train on identified data (stripped of personal identifiers) while still extracting and commercializing intellectual property value from conversations, essentially converting users' proprietary insights into unmarked training material.
  • Current service agreements with LLM providers lack NDAs, confidentiality clauses, or IP protection mechanisms, leaving organizations vulnerable to having their competitive insights distributed through open-source versions of models.

Topics

LLM training on user conversation dataIntellectual property protection and confidentialityModel training practices and lack of legal protectionsNovel research insights being absorbed into model trainingOpen-source model access to proprietary information

Transcript

[0:01] Should they trust any of these LLMs with their own knowledge, fearing it will be stolen into the core of the model? I've had experiences where we asked fairly new scientific questions and the model identified them as new insights. It's like, "Oh, I never thought about that ." "It would be interesting." Like, blah blah blah. And then, using a different account and requesting the same model later, or the next version, I encountered this. It's like, "Oh, well, you could do it like this." And it actually just [0:32] describes the exact thing we discussed in the previous version of the chat. Right now, these are just a few anecdotal cases. But I know the field we…

Full transcript available for MurmurCast members

Sign Up to Access

More from All-In Podcast

Get AI summaries like this delivered to your inbox daily

Get AI summaries delivered to your inbox

MurmurCast summarizes your YouTube channels, podcasts, and newsletters into one daily email digest.