The most-shared sentence from last week’s unsealed filings in the authors’ case against OpenAI comes from a Slack channel in July 2019. Sam McCandlish, then an OpenAI researcher, was explaining why he wanted to drop the word “LibGen” from a paper, and it turned out to predict its own future:
It showed up on Hacker News this weekend, under a headline about OpenAI fearing the optics of what might show up on Hacker News, with several hundred comments under it. The Authors Guild’s own press release puts the line under a heading that sums up its reading: “OpenAI Feared ‘Optics,’ Not the Law.”
The quote is real. I read it in the filing, and it says what the release says it does. What I wanted to know is narrower: what legal work does it do? The release is advocacy, which is fine, but it presents the documents as one story (“top execs knew”). The motion they were filed with asks the judge for five separate rulings, and knowledge is an element of none of them.
What the motion asks for
The release links two documents. One is the class plaintiffs’ memorandum of law (ECF 1982), 49 pages, served under seal on 4 September and filed publicly on 17 September. The other is the 162-page Rule 56.1 statement of facts (ECF 1987), where nearly every quote in the release comes from. The memorandum is a motion for partial summary judgment on liability. Its conclusion asks Judge Sidney Stein to hold:
- that the authors have made out a prima facie case of infringement;
- that OpenAI’s downloading from LibGen and Books3 was not fair use;
- that OpenAI’s copying of the books to train its models was not fair use;
- that OpenAI and Microsoft’s distribution of copies to each other and to contractors was not fair use;
- that Microsoft is vicariously liable for OpenAI’s infringement.
The motion covers 194 “Asserted Works” by the named authors, among them Baldacci, Franzen, Grisham, Martin and Picoult. It is not a motion about willfulness or damages.
Copyright infringement does not require a guilty mind. Fair use is judged on four statutory factors, none of which is “did the defendant think it was legal.” Vicarious liability turns on the right and ability to supervise plus a direct financial interest. So on the questions in front of the judge, the quotes the release put first mostly arrive as color. I checked this against the memorandum by tallying where each headline quote’s paragraphs in the fact statement are cited.
The count has a caveat. The argument section also refers back to the factual background as a whole (“supra”), and a tally of pinpoint citations misses that. It still shows where the lawyers expected the weight to fall. The two passages the argument keeps returning to are Jack Clark, then OpenAI’s policy director, writing in May 2020 that “our work on AI and Creativity is going to increasingly lead to us creating systems that substitute for the labor of [] people,” and Brent Hecht, a Microsoft applied-science director, writing that unlicensed training “has started a doom loop” because the product “threatens the economic foundations of its essential suppliers.” Those go to the first and fourth fair-use factors: purpose, and market harm. The optics line does not.
Downloading: knowledge is nearly surplus
On the downloading claim the authors hardly need the guilty-sounding quotes at all, because of how Bartz v. Anthropic came out. In June 2025 Judge William Alsup held that training Claude on books Anthropic had bought and scanned was fair use, and that building a library of pirated copies was a separate use that was not. The authors quote him: “Such piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded.” That rule turns on where the copies came from, not on what anyone at the company thought about it. When the memorandum does mention knowledge in its argument, the point is the price, not a guilty conscience: OpenAI “downloaded books for free while knowing it could have purchased or licensed copies precisely to avoid paying for them.”
OpenAI does not contest the copying itself. According to the fact statement, OpenAI employees torrented about 35 terabytes from LibGen between September 2019 and January 2020, which plaintiffs say matches the size of its whole 4.6-million-book fiction and nonfiction collections. The two training sets built from LibGen held roughly 570,000 books, and plaintiffs’ expert found 124 of the 194 asserted works in them. OpenAI’s defense to all of this is fair use, not denial. The memorandum says so outright: “Defendants do not dispute that they copied Plaintiffs’ works and millions of others.”
Training: the quotes that matter are about substitution
The training claim is the one the authors most need and are least sure to win. Judge Alsup held that training on lawfully bought books was fair use. Two days later, in Kadrey v. Meta, Judge Vince Chhabria ruled for Meta, but only because the plaintiffs had not built a record on what he called market dilution, and he said openly that such a record might have won. He wrote that “by training generative AI models with copyrighted works, companies are creating something that often will dramatically undermine the market for those works.” The authors in New York have read that as an instruction. Their argument on training is built around Clark, Hecht, and a record on AI-generated books entering the market, much of it redacted in the public copy.
I wrote about the same judge’s other half of this litigation a week ago, where the news publishers’ market-harm number turned out to rest on sealed Microsoft data. The authors’ version of the argument has the same weak point. Clark’s memo shows that OpenAI’s policy staff expected substitution in 2020. It does not measure substitution of these 194 books, and the passages that come closest to measuring it are blacked out in the public copy.
Where knowledge does count
The optics line and its neighbors will matter later, at damages. Under 17 U.S.C. § 504(c), statutory damages run from $750 to $30,000 per work. A court can raise the cap to $150,000 if the infringement was willful, or lower the floor to $200 if the defendant proves it “was not aware and had no reason to believe” it was infringing. Knowledge decides which band you are in.
For 194 works, the whole range tops out at $29.1 million. The number that would matter is the count of works, and that depends on a class, which none of the documents I read resolve. For scale, Bartz settled at $1.5 billion, which the court that approved it on 20 July described as roughly $3,000 per work, “four times the statutory minimum.” Apply that rate to the 570,000 books in OpenAI’s two LibGen sets and you get about $1.7 billion, though that is only an illustration. Not every one of those books is a registered US work, and a class would have to be certified first.
The willfulness fight has already been through one round in this case. On 24 November 2025, Magistrate Judge Ona Wang ruled that OpenAI had waived attorney-client privilege over why it deleted the LibGen datasets in 2022, in part because it was still denying willfulness: “For OpenAI to deny that it willfully infringed Class Plaintiffs’ copyrighted works is to argue that it acted in good faith.” She ordered the deletion communications produced and let plaintiffs depose OpenAI’s in-house lawyers. She also found that the crime-fraud exception did not apply. So the documents behind “excise-libgen” were pried loose because of willfulness, even though this motion does not ask for a willfulness ruling.
The names
The paper trail on names is the one part of this record a reader can check without a court login, because two of its steps are published papers.
In April 2019 a document prepared for a meeting with Bill Gates said the model had grown after OpenAI “added another ~11B words from Library Genesis (LibGen).” Three months later came the Slack thread above. The scaling-laws paper that came out in January 2020 tests on “a collection of publicly-available Internet Books.” The fact statement says those were the LibGen samples, and McCandlish later reminded colleagues, as they debated names for the next paper, that “we call it ‘Internet Books’ in the scaling laws paper.” Before the GPT-3 paper went out in May 2020, Dario Amodei asked in Slack: “Is it sketchy to call our corpuses ‘Books1’ and ‘Books2’ and not say what they are, particularly when in fact they are Libgen (which is a slightly sketchy source).” Ben Mann had told Clark the description was “deliberately vague since it’s libgen.” The GPT-3 paper went out with Books1 at 12 billion tokens and Books2 at 55 billion, each 8% of the training mix, described only as “two internet-based books corpora.”
For years that opacity was a research topic in its own right. A 2025 study of how books are valued in training data (Rowberry, in Convergence) could only note that Books2’s contents were unknown and that one lawsuit had guessed it was a shadow-library copy. Under oath, Mann said both sets were LibGen: “internally at OpenAI, Books1 and Books2 were known as LibGen1 and LibGen2.” The training-file names in the fact statement are blunter still: one version of Books1 used for GPT-3.5 was “libgen1-dedup-clean-csam-filt.”
One more detail sits in the background. Several of the people in this story, including Amodei, McCandlish, Mann, Clark, Tom Brown and Jared Kaplan, the scaling-laws paper’s first author, left OpenAI and are listed among the founders of Anthropic, the company that paid $1.5 billion in Bartz for doing much the same with LibGen. It has no legal bearing on this case. What it shows is that in 2019 and 2020, people who went on to build a company around AI risk treated the provenance of books as an optics problem, and the courts have since treated it as a legal one. It is also the second time in a week that a question of state of mind has turned out to matter less than which provision is being applied. Anthropic’s two Pentagon cases split on exactly that: motive decided one statute and was irrelevant under the other.
What to watch
OpenAI’s reply to these paragraphs is not yet public as far as I can find, and a fact statement is one side’s list of what it says is undisputed. Many of these facts will be disputed. When Judge Stein rules, I would read the opinion for the word “willful” and expect not to find it, at least not as a holding. What I would look for is whether he adopts Alsup’s split, with downloading as a separate use judged on its own terms, and whether he thinks the substitution record is stronger than the one Chhabria found missing in Kadrey. The optics message will probably get quoted in the opinion’s background. The ruling itself will turn on the Napster line of cases and on the question of substitution.
References
- Authors Guild (2026). Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI: Top Execs Knew Their Mass Book Piracy Was Illegal And Would Put Authors Out of Work. Press release, 21 September 2026.
- United States District Court, S.D.N.Y. In re OpenAI, Inc. Copyright Infringement Litigation, No. 1:25-md-03143-SHS-OTW. Class Plaintiffs’ Memorandum of Law in Support of Motion for Partial Summary Judgment, ECF 1982. Served 4 September 2026, filed publicly 17 September 2026. Citation tally in figure 2 is the writer’s own count.
- United States District Court, S.D.N.Y. In re OpenAI, Inc. Copyright Infringement Litigation, No. 1:25-md-03143-SHS-OTW. Class Plaintiffs’ Corrected Rule 56.1 Statement of Undisputed Material Facts, ECF 1987. Filed 17 September 2026; quoted paragraphs 188, 193–194, 269–282, 312–315, 341–343.
- United States District Court, S.D.N.Y. In re OpenAI, Inc. Copyright Infringement Litigation, No. 25-md-3143. Opinion & Order re: OpenAI’s Deletion of Books1 and Books2 Datasets and Privilege Rulings. Wang, M.J., 24 November 2025.
- 17 U.S.C. § 504, Remedies for infringement: Damages and profits. Legal Information Institute, accessed 27 September 2026.
- Authors Guild (2026). Court Grants Final Approval of $1.5 Billion Anthropic Copyright Settlement. 21 July 2026, reporting the 20 July 2026 order in Bartz v. Anthropic.
- Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007 (N.D. Cal. 2025) (Alsup, J., 23 June 2025). As quoted in reference 2.
- Kadrey v. Meta Platforms, Inc., 788 F. Supp. 3d 1026 (N.D. Cal. 2025) (Chhabria, J., 25 June 2025). As quoted in reference 2.
- Kaplan, J., McCandlish, S., et al. (2020). Scaling Laws for Neural Language Models. arXiv:2001.08361, 23 January 2020.
- Brown, T. B., Mann, B., et al. (2020). Language Models are Few-Shot Learners. arXiv:2005.14165, 28 May 2020. Table 2.2.
- Rowberry, S. (2025). The value of books in the age of generative AI training data. Convergence, 31(6), 1935–1950.
- Wikipedia. Anthropic, founders list. Accessed 27 September 2026.
- Hacker News (2026). OpenAI Feared “Optics” of what might appear on Hacker News. Discussion thread, accessed 27 September 2026.