OpenAI staff called the pirated book source behind early ChatGPT training ‘sketchy AF’ in newly unsealed internal messages

by | Sep 20, 2026 | Latest E-commerce News & Updates

OpenAI staff knew the books behind early ChatGPT models came from Library Genesis, a piracy site, and discussed hiding it, according to a filing unsealed Thursday in the copyright case John Grisham, David Baldacci, and other authors brought against OpenAI and Microsoft. David Lansky, OpenAI’s general counsel at the time, put LibGen forward as a data source in 2019, and GPT-3 engineer Tom Brown and technical staffer Ben Mann called it “sketchy AF.” Anthropic CEO Dario Amodei, who worked as a senior researcher at OpenAI at the time, asked in Slack whether naming the corpora Books1 and Books2 without saying what they were counted as sketchy, and Mann wrote that one description was “deliberately vague since it’s libgen.” OpenAI pulled the data in 2022 after Bob McGrew said that would be “very valuable for legal reasons.”

Paul Drecksler is the founder and editor of Shopifreaks, covering the most important stories in e-commerce.

Companies: OpenAI

Never miss important e-commerce news

Our weekly newsletter is read religiously by 20,000+ e-commerce professionals.

Loading...