TCG082: AI News Roundtable – Copyrights, AI Watermarks, and the Open Weight Debate
This news roundup covers three major AI industry developments: Anthropic's introduction of watermarks for AI-generated content, a $1.5 billion copyright settlement over illegal book acquisition, and the emergence of open-weight models like Kimi K3 as alternatives to closed commercial systems. The hosts discuss the tensions between these moves and broader questions about fair use, model training practices, and industry hypocrisy.
Summary
William and Yvonne discuss a concentrated period of AI industry news centered on Anthropic. The episode begins with Anthropic's announcement of two watermarking mechanisms: imperceptible text-based watermarks embedded at the token level in Claude's outputs, and C2PA standard digital signatures on generated files. While Yvonne views watermarks as a necessary step toward maturity in the industry—particularly for creative works and code—both hosts note that Anthropic's own caveats undermine the watermark's conclusiveness: a detected watermark doesn't prove AI generation, and its absence proves nothing about human authorship. They note that watermark-stripping tools already exist and can be readily deployed in CI/CD pipelines.
The second major topic is Anthropic's $1.5 billion copyright settlement, the largest in American history, covering approximately 500,000 works. The court documents reveal that Anthropic purchased bulk used books, destructively scanned them using hydraulic spine-cutters and industrial scanners, and recycled the paper—an initiative called Project Panorama. Yvonne argues this is relatively uncontroversial, comparing it to typical flea market book acquisition and noting that format-shifting of physical books to digital copies constitutes fair use. However, William emphasizes that the problematic acquisition method—torrenting millions of pirated books from shadow libraries—was what violated copyright, not the subsequent training. The court ruled that acquiring books legally and scanning them constitutes transformative fair use, creating what William calls a "clean path" forward.
The third topic examines the contradiction between Anthropic's actions and public identity. Anthropic has accused Chinese AI labs (Moonshot, DeepSeek) of conducting "industrial skill distillation attacks" by creating fake accounts to extract Claude's outputs for training their own models. William points out the moral inconsistency: Anthropic acquired 7 million pirated books and argued learning from them was transformative, yet characterizes Moonshot's use of Anthropic's outputs as an attack. He frames this as "distillation for me but not for thee." This contradiction, combined with Anthropic's safety-forward public branding versus its actual practices, creates what William describes as a "brand tension."
The episode concludes by discussing Kimi K3, an open-weight model from Moonshot AI with 2.8+ trillion parameters, released alongside open-sourced attention kernels and agent tooling. This prompted Jensen Huang's open letter signed by hundreds of AI industry figures (notably excluding Anthropic) arguing for open-weight models as essential to American AI leadership, drawing parallels to the open-source software movement. Yvonne expresses strong support for open-weight models as necessary counterbalance to closed commercial development, citing how the open-source movement created net economic good without damage. She argues that both open and closed approaches have value and can feed each other productively. William notes the tension between concentrated CapEx recovery requirements and avoiding technological monopolies.
About this episode
William Collins and Eyvonne Sharp dig into the latest AI headlines, from the largest copyright settlement in American history to stolen AI models and invisible watermarks on Claude output. Plus, they discuss why so many companies have rallied around NVIDIA’s support for open weight AI models. Our hosts also examine the biggest questions arising from<a class="excerpt-read-more" href="https://packetpushers.net/podcasts/the-cloud-gambit/tcg082-ai-news-roundtable-copyrights-ai-watermarks-and-the-open-weight-debate/" title="ReadTCG082: AI News Roundtable – Copyrights, AI Watermarks, and the Open Weight Debate">... Read more »</a>
Key Insights
- Anthropic's watermarking solution is self-undermining because their own documentation states that detected watermarks are not conclusive proof of AI generation and absent watermarks prove nothing, making the watermark merely advisory in multiple directions rather than definitive.
- The court distinguished between how content is acquired and how it is used: illegal torrent acquisition violated copyright, but legal acquisition and format-shifting of physical books was ruled transformative fair use, creating a clear legal pathway forward for Anthropic.
- Anthropic downloaded 7 million books through shadow library torrents and successfully argued their use was transformative, yet accused Chinese labs of conducting attacks when those labs used Anthropic's own outputs for training—applying contradictory ethical standards to identical activities.
- Yvonne argues that destructively scanning bulk used books through industrial processes is relatively uncontroversial and preferable to illegal torrenting, comparing the practice to flea market book acquisition rather than symbolic book burning.
- William contends that the AI industry has taken a 'horrendously low EQ approach' to public communication, overstating AI's impact on employment and creating unnecessary social tension that didn't need to exist.
- Open-weight models like Kimi K3 represent a counterbalance to closed commercial development similar to how open-source software prevented software monopolies while creating net economic good without damage to the industry.
- The watermarking approach can detect patterns in grammatically precise writing as AI-generated, which problematically penalizes good writing and has weakened AI detection conversations in academic spaces.
- Anthropic's brand identity as a safety-forward laboratory conflicts with its practice of acquiring pirated content and applying double standards to competitive practices, creating tension between public principles and actual behavior.
Topics
Transcript
. In about three weeks time this summer, one AI company managed to do all of the following. It got hit with the largest copyright settlement in American history. It had internal documents unsealed showing it bought millions of used books, sliced off the spines with a hydraulic cutter, and scanned them in. But the paper was recycled, so there is that. It accused Chinese labs of stealing its models when those labs asked questions and used just the output to train their models. And then this morning, and this is the juicy part, it announced it's going to start hiding invisible watermarks in everything it writes, including your code. So the rest of us can tell what came from…
Full transcript available for MurmurCast members
Sign Up to AccessMore from The Everything Feed - All Packet Pushers Pods
TNO071: The Network Team Is Drowning. Is AI the Life Raft? (Sponsored)
Rekha Shenoy and Irfan Kimji from Backbox discuss how the exponential growth of vulnerabilities (49,000 CVEs annually) has made manual network operations unsustainable, and how AI-powered automation can help network teams manage patches and security updates at scale while maintaining human control and oversight.
HN840: How to Make a Technology Buying Decision
Sean Morgan, a research director at Deloro Group, discusses how technology buying decisions should extend beyond engineering specifications to include business alignment, ROI calculations, and understanding total cost of ownership. Engineers must shift from viewing IT as a cost center to positioning it as a business enabler by connecting technical decisions to revenue impact and organizational objectives.
IPB207: Flying Blind: Monitoring Might Not See IPv6
The IPv6 Buzz hosts discuss critical gaps in IPv6 monitoring across enterprise networks, highlighting that many monitoring platforms lack IPv6 awareness, vendor parity, and advanced analytical capabilities. They emphasize that while basic IPv6 data ingestion has improved, sophisticated features like cross-protocol event correlation, extension header analysis, and device identity tracking remain significant industry challenges.
N4N063: Link Layer Discovery Protocol
Link Layer Discovery Protocol (LLDP) is a standardized Layer 2 protocol that enables network devices to announce information about themselves to directly connected neighbors, facilitating network topology discovery and device identification in multi-vendor environments. The protocol uses Ethernet frames with special multicast destination MAC addresses to ensure frames don't propagate beyond immediate neighbors, and includes mandatory TLVs (Type-Length-Values) like chassis ID, port ID, and TTL alongside optional ones for extended information.
TCG083: Superintelligence for Everyone: Who Actually Holds the Power?
Three technology experts discuss Mark Zuckerberg's manifesto on distributed superintelligence, examining whether his promises of universal access and individual empowerment align with infrastructure realities. They conclude that while decentralized AI is theoretically safer than centralized control, the manifesto fails to account for human complexity, existing inequalities, and the enormous capital requirements that will likely concentrate power rather than distribute it.