OpenAI staff knew the books behind early ChatGPT models came from Library Genesis, a piracy site, and discussed hiding it, according to a filing unsealed Thursday in the copyright case John Grisham, David Baldacci, and other authors brought against OpenAI and Microsoft. David Lansky, OpenAI’s general counsel at the time, put LibGen forward as a data source in 2019, and GPT-3 engineer Tom Brown and technical staffer Ben Mann called it “sketchy AF.” Anthropic CEO Dario Amodei, who worked as a senior researcher at OpenAI at the time, asked in Slack whether naming the corpora Books1 and Books2 without saying what they were counted as sketchy, and Mann wrote that one description was “deliberately vague since it’s libgen.” OpenAI pulled the data in 2022 after Bob McGrew said that would be “very valuable for legal reasons.”






