OpenAI Worried About the Optics. The Authors' Brief Leans on Something Else.

The most-shared sentence from last week’s unsealed filings in the authors’ case against OpenAI comes from a Slack channel in July 2019. Sam McCandlish, then an OpenAI researcher, was explaining why he wanted to drop the word “LibGen” from a paper, and it turned out to predict its own future:

Paragraphs 341 to 343 of the class plaintiffs' Rule 56.1 statement. Paragraph 341 quotes a July 19, 2019 Slack exchange in which Sam McCandlish proposes removing all mentions of LibGen from a scaling paper, Dario Amodei calls LibGen as a training set a bit sketchier, and McCandlish says he was worried about optics. Paragraph 343 says the published scaling-laws paper described the LibGen data as Internet Books.
Paragraph 341 is the one that went viral. Paragraph 343 is the one you can check yourself, by opening the published paper. Image: cropped from page 56 of 162, Class Plaintiffs' Corrected Rule 56.1 Statement of Undisputed Material Facts, ECF 1987, In re OpenAI, Inc. Copyright Infringement Litigation, No. 1:25-md-03143 (S.D.N.Y.), filed 17 September 2026, via the Authors Guild. A public court filing, reproduced for commentary.

It showed up on Hacker News this weekend, under a headline about OpenAI fearing the optics of what might show up on Hacker News, with several hundred comments under it. The Authors Guild’s own press release puts the line under a heading that sums up its reading: “OpenAI Feared ‘Optics,’ Not the Law.”

The quote is real. I read it in the filing, and it says what the release says it does. What I wanted to know is narrower: what legal work does it do? The release is advocacy, which is fine, but it presents the documents as one story (“top execs knew”). The motion they were filed with asks the judge for five separate rulings, and knowledge is an element of none of them.

What the motion asks for

The release links two documents. One is the class plaintiffs’ memorandum of law (ECF 1982), 49 pages, served under seal on 4 September and filed publicly on 17 September. The other is the 162-page Rule 56.1 statement of facts (ECF 1987), where nearly every quote in the release comes from. The memorandum is a motion for partial summary judgment on liability. Its conclusion asks Judge Sidney Stein to hold:

  1. that the authors have made out a prima facie case of infringement;
  2. that OpenAI’s downloading from LibGen and Books3 was not fair use;
  3. that OpenAI’s copying of the books to train its models was not fair use;
  4. that OpenAI and Microsoft’s distribution of copies to each other and to contractors was not fair use;
  5. that Microsoft is vicariously liable for OpenAI’s infringement.

The motion covers 194 “Asserted Works” by the named authors, among them Baldacci, Franzen, Grisham, Martin and Picoult. It is not a motion about willfulness or damages.

Copyright infringement does not require a guilty mind. Fair use is judged on four statutory factors, none of which is “did the defendant think it was legal.” Vicarious liability turns on the right and ability to supervise plus a direct financial interest. So on the questions in front of the judge, the quotes the release put first mostly arrive as color. I checked this against the memorandum by tallying where each headline quote’s paragraphs in the fact statement are cited.

Where the headline quotes are cited in the authors' summary-judgment memorandum A dot count of citations to the Rule 56.1 statement in the 49-page memorandum, ECF 1982, split between the introduction and factual background and the argument section. Jack Clark's substitution memo, paragraphs 749 to 750: 7 citations in the facts, 5 in the argument. Brent Hecht's doom loop, paragraph 765: 2 and 3. Gogineni's tweets about finishing George R. R. Martin's series, paragraphs 727 to 731: 4 and 1, the one being a sweeping cite to paragraphs 706 to 916. Altman and Nadella agreeing piracy is illegal, paragraphs 370 to 384: 3 and 0. We trained GPT-3 on pirated stuff, paragraph 344: 3 and 0. The Gates meeting disclosure, paragraphs 182 to 191: 2 and 0. Renaming LibGen to Books1 and Books2, paragraphs 269 to 282: 1 and 0. The excise-libgen deletion, paragraphs 312 to 316: 2 and 0. The optics message, paragraph 341: 1 and 0, and that one is inside a string cite of paragraphs 337 to 367. CITATIONS TO THE FACT STATEMENT · CLASS PLAINTIFFS' MEMORANDUM, ECF 1982 PRE-ARGUMENT · PP. 1–17 ARGUMENT · PP. 18–41 Clark: systems that "substitute" Hecht: a "doom loop" Gogineni: "GPT-5 will autocomplete" Altman, Nadella: piracy is illegal "We trained GPT-3 on pirated stuff!" LibGen disclosed at the Gates meeting LibGen renamed "Books1", "Books2" "excise-libgen" / Project Clear "I was just worried about optics" ¶¶749–50 ¶765 ¶¶727–31 ¶¶370–84 ¶344 ¶¶182–91 ¶¶269–82 ¶¶312–16 ¶341 5 3 1 000000 One dot per citation, my count; ranges expanded, so broad cites like ¶¶706–916 count. ¶341 is cited only within ¶¶337–367.
The argument section leans on two witnesses, and both are talking about substitution. The knowledge and concealment material stays in the story the brief tells, not in the legal analysis.

The count has a caveat. The argument section also refers back to the factual background as a whole (“supra”), and a tally of pinpoint citations misses that. It still shows where the lawyers expected the weight to fall. The two passages the argument keeps returning to are Jack Clark, then OpenAI’s policy director, writing in May 2020 that “our work on AI and Creativity is going to increasingly lead to us creating systems that substitute for the labor of [] people,” and Brent Hecht, a Microsoft applied-science director, writing that unlicensed training “has started a doom loop” because the product “threatens the economic foundations of its essential suppliers.” Those go to the first and fourth fair-use factors: purpose, and market harm. The optics line does not.

Downloading: knowledge is nearly surplus

On the downloading claim the authors hardly need the guilty-sounding quotes at all, because of how Bartz v. Anthropic came out. In June 2025 Judge William Alsup held that training Claude on books Anthropic had bought and scanned was fair use, and that building a library of pirated copies was a separate use that was not. The authors quote him: “Such piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded.” That rule turns on where the copies came from, not on what anyone at the company thought about it. When the memorandum does mention knowledge in its argument, the point is the price, not a guilty conscience: OpenAI “downloaded books for free while knowing it could have purchased or licensed copies precisely to avoid paying for them.”

OpenAI does not contest the copying itself. According to the fact statement, OpenAI employees torrented about 35 terabytes from LibGen between September 2019 and January 2020, which plaintiffs say matches the size of its whole 4.6-million-book fiction and nonfiction collections. The two training sets built from LibGen held roughly 570,000 books, and plaintiffs’ expert found 124 of the 194 asserted works in them. OpenAI’s defense to all of this is fair use, not denial. The memorandum says so outright: “Defendants do not dispute that they copied Plaintiffs’ works and millions of others.”

Training: the quotes that matter are about substitution

The training claim is the one the authors most need and are least sure to win. Judge Alsup held that training on lawfully bought books was fair use. Two days later, in Kadrey v. Meta, Judge Vince Chhabria ruled for Meta, but only because the plaintiffs had not built a record on what he called market dilution, and he said openly that such a record might have won. He wrote that “by training generative AI models with copyrighted works, companies are creating something that often will dramatically undermine the market for those works.” The authors in New York have read that as an instruction. Their argument on training is built around Clark, Hecht, and a record on AI-generated books entering the market, much of it redacted in the public copy.

I wrote about the same judge’s other half of this litigation a week ago, where the news publishers’ market-harm number turned out to rest on sealed Microsoft data. The authors’ version of the argument has the same weak point. Clark’s memo shows that OpenAI’s policy staff expected substitution in 2020. It does not measure substitution of these 194 books, and the passages that come closest to measuring it are blacked out in the public copy.

Where knowledge does count

The optics line and its neighbors will matter later, at damages. Under 17 U.S.C. § 504(c), statutory damages run from $750 to $30,000 per work. A court can raise the cap to $150,000 if the infringement was willful, or lower the floor to $200 if the defendant proves it “was not aware and had no reason to believe” it was infringing. Knowledge decides which band you are in.

Statutory damages per work under 17 U.S.C. 504(c), on a log scale Five per-work amounts on a logarithmic axis from 100 to 200,000 dollars. The innocent-infringement floor is 200 dollars, which is 38,800 dollars across the 194 works in the motion. The ordinary minimum is 750 dollars, or 145,500 dollars across 194 works. The Bartz v. Anthropic settlement paid about 3,000 dollars per work, or about 582,000 dollars at 194 works. The ordinary maximum is 30,000 dollars, or 5.82 million. The willful maximum is 150,000 dollars, or 29.1 million. The band from 30,000 to 150,000 dollars is reachable only with a willfulness finding. DOLLARS PER INFRINGED WORK · LOG SCALE · 17 U.S.C. § 504(c) WILLFUL ONLY $100$1k$10k$100k $200 — innocent-infringement floor, § 504(c)(2) $750 — ordinary minimum, § 504(c)(1) $30,000 — ordinary maximum, § 504(c)(1) $150,000 — willful maximum, § 504(c)(2) About $3,000 — Bartz v. Anthropic settlement, per work, final approval 20 July 2026 Innocentfloor $200 Ordinarymin $750 Bartz deal~$3,000 Ordinarymax $30k Willfulmax $150k × 194 ASSERTED WORKS $38.8k $145.5k ~$582k $5.82M $29.1M Statutory figures from the statute; the Bartz figure is a settlement average, not an award. Totals are my arithmetic.
At the scale of this motion, even the willful ceiling is small money for OpenAI. The exposure only becomes large if the count of works grows toward the hundreds of thousands in the LibGen sets.

For 194 works, the whole range tops out at $29.1 million. The number that would matter is the count of works, and that depends on a class, which none of the documents I read resolve. For scale, Bartz settled at $1.5 billion, which the court that approved it on 20 July described as roughly $3,000 per work, “four times the statutory minimum.” Apply that rate to the 570,000 books in OpenAI’s two LibGen sets and you get about $1.7 billion, though that is only an illustration. Not every one of those books is a registered US work, and a class would have to be certified first.

The willfulness fight has already been through one round in this case. On 24 November 2025, Magistrate Judge Ona Wang ruled that OpenAI had waived attorney-client privilege over why it deleted the LibGen datasets in 2022, in part because it was still denying willfulness: “For OpenAI to deny that it willfully infringed Class Plaintiffs’ copyrighted works is to argue that it acted in good faith.” She ordered the deletion communications produced and let plaintiffs depose OpenAI’s in-house lawyers. She also found that the crime-fraud exception did not apply. So the documents behind “excise-libgen” were pried loose because of willfulness, even though this motion does not ask for a willfulness ruling.

The names

The paper trail on names is the one part of this record a reader can check without a court login, because two of its steps are published papers.

What OpenAI called its LibGen book sets, April 2019 to September 2026 A timeline drawn from the plaintiffs' Rule 56.1 statement. April 2019: a document prepared for a meeting with Bill Gates says OpenAI added about 11 billion words from Library Genesis. July 2019: a Slack message proposes removing all mentions of LibGen from a scaling paper. January 2020: the published scaling-laws paper calls the data Internet Books. May 2020: the GPT-3 paper calls the two LibGen sets Books1 and Books2. June 2022: a Slack channel named excise-libgen is renamed project-clear. September 2026: the statement is filed publicly. Spacing is proportional to time up to 2022; the axis is broken before 2026. THE NAME ON THE BOOKS · CLASS PLAINTIFFS' RULE 56.1 STATEMENT (ECF 1987) April 2019 — Gates meeting document, SUF ¶188 19 July 2019 — Slack, SUF ¶341 January 2020 — "Internet Books", SUF ¶343 May 2020 — "Books1" and "Books2", SUF ¶¶276–279 15 June 2022 — excise-libgen renamed project-clear, SUF ¶¶312–315 17 September 2026 — public filing of ECF 1987 Apr 2019 Jul 2019 Jan 2020 May 2020 Jun 2022 Sep 2026 Gates doc: "Library Genesis" Slack: remove "all mentions of LibGen" Scaling-laws paper: "Internet Books" GPT-3 paper: "Books1" and "Books2" "excise-libgen" channel renamed "project-clear" Statement filed publicly Spacing is proportional to time from April 2019 to June 2022; the axis is broken before 2026. Paragraph citations in the tooltips.
Same books, four names. The two in bold are the published papers. The rest come from internal documents as the plaintiffs quote them, and OpenAI has not yet had its turn to answer.

In April 2019 a document prepared for a meeting with Bill Gates said the model had grown after OpenAI “added another ~11B words from Library Genesis (LibGen).” Three months later came the Slack thread above. The scaling-laws paper that came out in January 2020 tests on “a collection of publicly-available Internet Books.” The fact statement says those were the LibGen samples, and McCandlish later reminded colleagues, as they debated names for the next paper, that “we call it ‘Internet Books’ in the scaling laws paper.” Before the GPT-3 paper went out in May 2020, Dario Amodei asked in Slack: “Is it sketchy to call our corpuses ‘Books1’ and ‘Books2’ and not say what they are, particularly when in fact they are Libgen (which is a slightly sketchy source).” Ben Mann had told Clark the description was “deliberately vague since it’s libgen.” The GPT-3 paper went out with Books1 at 12 billion tokens and Books2 at 55 billion, each 8% of the training mix, described only as “two internet-based books corpora.”

For years that opacity was a research topic in its own right. A 2025 study of how books are valued in training data (Rowberry, in Convergence) could only note that Books2’s contents were unknown and that one lawsuit had guessed it was a shadow-library copy. Under oath, Mann said both sets were LibGen: “internally at OpenAI, Books1 and Books2 were known as LibGen1 and LibGen2.” The training-file names in the fact statement are blunter still: one version of Books1 used for GPT-3.5 was “libgen1-dedup-clean-csam-filt.”

One more detail sits in the background. Several of the people in this story, including Amodei, McCandlish, Mann, Clark, Tom Brown and Jared Kaplan, the scaling-laws paper’s first author, left OpenAI and are listed among the founders of Anthropic, the company that paid $1.5 billion in Bartz for doing much the same with LibGen. It has no legal bearing on this case. What it shows is that in 2019 and 2020, people who went on to build a company around AI risk treated the provenance of books as an optics problem, and the courts have since treated it as a legal one. It is also the second time in a week that a question of state of mind has turned out to matter less than which provision is being applied. Anthropic’s two Pentagon cases split on exactly that: motive decided one statute and was irrelevant under the other.

What to watch

OpenAI’s reply to these paragraphs is not yet public as far as I can find, and a fact statement is one side’s list of what it says is undisputed. Many of these facts will be disputed. When Judge Stein rules, I would read the opinion for the word “willful” and expect not to find it, at least not as a holding. What I would look for is whether he adopts Alsup’s split, with downloading as a separate use judged on its own terms, and whether he thinks the substitution record is stronger than the one Chhabria found missing in Kadrey. The optics message will probably get quoted in the opinion’s background. The ruling itself will turn on the Napster line of cases and on the question of substitution.

References

  1. Authors Guild (2026). Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI: Top Execs Knew Their Mass Book Piracy Was Illegal And Would Put Authors Out of Work. Press release, 21 September 2026.
  2. United States District Court, S.D.N.Y. In re OpenAI, Inc. Copyright Infringement Litigation, No. 1:25-md-03143-SHS-OTW. Class Plaintiffs’ Memorandum of Law in Support of Motion for Partial Summary Judgment, ECF 1982. Served 4 September 2026, filed publicly 17 September 2026. Citation tally in figure 2 is the writer’s own count.
  3. United States District Court, S.D.N.Y. In re OpenAI, Inc. Copyright Infringement Litigation, No. 1:25-md-03143-SHS-OTW. Class Plaintiffs’ Corrected Rule 56.1 Statement of Undisputed Material Facts, ECF 1987. Filed 17 September 2026; quoted paragraphs 188, 193–194, 269–282, 312–315, 341–343.
  4. United States District Court, S.D.N.Y. In re OpenAI, Inc. Copyright Infringement Litigation, No. 25-md-3143. Opinion & Order re: OpenAI’s Deletion of Books1 and Books2 Datasets and Privilege Rulings. Wang, M.J., 24 November 2025.
  5. 17 U.S.C. § 504, Remedies for infringement: Damages and profits. Legal Information Institute, accessed 27 September 2026.
  6. Authors Guild (2026). Court Grants Final Approval of $1.5 Billion Anthropic Copyright Settlement. 21 July 2026, reporting the 20 July 2026 order in Bartz v. Anthropic.
  7. Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007 (N.D. Cal. 2025) (Alsup, J., 23 June 2025). As quoted in reference 2.
  8. Kadrey v. Meta Platforms, Inc., 788 F. Supp. 3d 1026 (N.D. Cal. 2025) (Chhabria, J., 25 June 2025). As quoted in reference 2.
  9. Kaplan, J., McCandlish, S., et al. (2020). Scaling Laws for Neural Language Models. arXiv:2001.08361, 23 January 2020.
  10. Brown, T. B., Mann, B., et al. (2020). Language Models are Few-Shot Learners. arXiv:2005.14165, 28 May 2020. Table 2.2.
  11. Rowberry, S. (2025). The value of books in the age of generative AI training data. Convergence, 31(6), 1935–1950.
  12. Wikipedia. Anthropic, founders list. Accessed 27 September 2026.
  13. Hacker News (2026). OpenAI Feared “Optics” of what might appear on Hacker News. Discussion thread, accessed 27 September 2026.