Events
Colloquium
Closing the book on memorization, language models, and copyright
|
Speaker: A. Feder Cooper (Yale) Assistant Professor of Computer Science Yale University Wednesday, September 30, 2026 12:00PM - 1:00PM Lunch 11:30am in 1307
Talk 12:00pm in 1327 Location: Yale Institute for Foundations of Data Science, Kline Tower 13th Floor, Room 1327, New Haven, CT 06511 and via Webcast: https://yale.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=328534ec-7916-4cbb-becb-b4cf011ed693 |
Abstract: Copyright disputes over generative AI increasingly turn on claims that models “memorize” copyrighted works in their training data, and that this memorized content can later be “extracted” (near-)verbatim in model outputs. When a language model reproduces its training data in an output, that output is a “copy” in the sense copyright law cares about. But the fact that a model can reproduce a work from its training data also tells us something about what is encoded inside the model itself. This raises a harder (and potentially very expensive) question: is the model a “copy” of the training data it has memorized? The answer turns out to be very complicated, and hinges on surprisingly deep technical details at the intersection of machine learning and copyright law.
For the last several years, my colleagues and I have tried to refine the answer to this question in various ways. On the machine learning side, we’ve developed new ways to measure memorization, sharpening what exactly it means when one says that a language model has “memorized” its training data. And on the legal side, I’ve worked with copyright scholars to pin down which technical facts about generative AI systems might actually matter for the law, even as the underlying science has continued to evolve underneath us.
In this talk, I will synthesize some of what we’ve learned, and where I think the frontier on these questions currently stands. To do so, I will center the talk on a large-scale study of memorization and extraction of copyrighted books in open-weight language models. In this work, we found that the extent of memorization varies substantially by model and by book. While most models don’t seem to memorize most books, there are striking exceptions. In extreme cases, memorization is extensive enough that one can deterministically generate near-pristine copies of entire books using only their opening words as the initial prompt. I will also briefly touch on how these findings translate to production systems, such as Claude and Gemini.
Along the way, I will discuss the measurement methods needed to make careful, rigorous claims about memorization and extraction: what counts as evidence that a model has memorized something, how extractability in outputs differs from memorization in the model, and why seemingly mundane methodological choices can mean the difference between valid and invalid evidence of memorization.
Speaker Bio: I’m an Assistant Professor of Computer Science at Yale University, where I’m also affiliated with the Information Society Project at Yale Law School, the Center for Algorithms, Data, and Market Design, and the Institute for Foundations of Data Science. I’m also an Affiliate Researcher at Stanford, working with Percy Liang, Dan Ho, and Mark Lemley, and a Faculty Associate at the Berkman Klein Center for Internet & Society at Harvard University.
My research has received spotlights, orals, and best-paper accolades at top AI/ML and computing venues, including NeurIPS, ICML, AAAI, and AIES. Work on copyright and Generative AI has been lauded as “landmark” work among scholars and the popular press. My research has been covered in the media at outlets such as The Atlantic, The Washington Post, Bloomberg News, 404 Media, and Wired.
Add To: Google Calendar | Outlook | iCal File
- Colloquium
Submit an Event
Interested in creating your own event, or have an event to share? Please fill the form if you’d like to send us an event you’d like to have added to the calendar.
