On 4 September the news publishers suing OpenAI and Microsoft filed a 92-page summary judgment brief in the Southern District of New York. The copy that anyone can download from the public court archive has a sentence on its tenth page that reads, in full, like this:
Defendants’ own data show this [blank] in motion, with Microsoft recording [blank] drops in click-through rates for The Times and DNP’s domains, and [blank] for ZD’s domains, for Microsoft’s Copilot “answer engine” compared to traditional Bing search. SF1536.
On 17 September the blanks were filled in. By that evening the number in the headlines was 93%. TechCrunch reported that Copilot “caused click-through rates for The New York Times’ domain to drop as much as 93%,” and that is the figure that travelled. Bloomberg Law printed the full version: “an 83-93% drop in click-through rates for Times’ and Daily News’ content, and 51-94% for Ziff Davis’ domains.”
So the most quoted number in the case is the high end of one range, and it is not even the highest number in the sentence. That is a small thing. What it tells you about the rest of this week’s coverage is not.
TechCrunch, to its credit, said the important part out loud: “much of the new information comes from The Times’ own brief, not the underlying exhibits, which remain sealed. The quotes below are presented without their original context.” That caveat deserves more weight than it got, because the 93% sits at the end of a chain with three links, and only the last one is public.
The first link is whatever Microsoft actually recorded, whether a dataset or a slide. The second is paragraph 1536 of the plaintiffs’ statement of facts, which is where a lawyer summarised that record. The third is the sentence in the brief, which is a lawyer summarising the summary. Nobody outside the case has seen the first link, and I could not find the statement of facts among the public filings.
That chain leaves out most of what you would need to read the number. I do not know the baseline click-through rate on Bing that the drop is measured against. I do not know the period, the query set, or whether “click-through” in an answer engine means the same thing as it does on a results page, where a user sees ten links and picks one. In Copilot the unit could plausibly be clicks per cited source, or per answer, or per session. A drop from 10% to 0.7% and a drop from 0.3% to 0.021% are both 93%, and they describe very different businesses. The number may well be accurate. Its denominator is sealed.
The same problem applies, more sharply, to the words. This week’s coverage treated five phrases as roughly the same kind of evidence. Put each one back against the public copy of the 4 September brief and they come apart.
Take “doom loop.” It is Brent Hecht’s phrase, and it is real: it appears in the publicly filed brief for the book authors, dated 17 September, which says Hecht, Microsoft’s Director of Applied Science, “wrote that Microsoft’s failure to license data for training ‘has started a doom loop’ because the LLM ‘end product threatens the economic foundations of its essential suppliers.’” In that document the loop is about licensing. In the news brief it does a different job. The one place it is legible in the public copy, on page 73, it belongs to an unnamed “Microsoft executive” who “has said [it] would imperil the future of AI itself.” On page 10 it is blacked out, but Bloomberg Law’s account of the unsealed version fills in the gap: the doom loop “is already in motion and demonstrated by drops in click-through rates for news content, according to the publishers’ motion.” So in one brief the loop is about licensing, and in the other it is the click data. TechCrunch goes a step further and says Hecht’s January 2024 presentation itself “describes the decline as a ‘doom loop.’” That may be right. But in the public record the sentence joining the phrase to the numbers is the lawyers’ (“Defendants’ own data show this [blank] in motion”), and from outside I cannot tell whether the deck and the data are the same document. Neither could anyone writing on Thursday night.
“Existential threat” is odder. TechCrunch reports that Nick Turley, OpenAI’s head of ChatGPT, wrote that publishers face an “existential threat” from products that are “largely substitutive.” In the public copy of the news brief, Turley’s words are blacked out. Bloomberg Law, working from the unsealed version, quotes him as writing “[o]ur products are largely substitutive, period,” and does not use the word existential. Where the phrase does appear in the public record is the authors’ brief, twice, both times in the lawyers’ own voice. Page 7: “OpenAI’s GPT models pose an existential threat to those who write and publish books.” Turley may well have written it too. The unredacted news brief may say so. But the only place I can see those two words on the record, they are advocacy, not admission.
“The largest theft of labor in human history” comes, per TechCrunch, from a January 2023 memo by Hecht, whom TechCrunch’s report calls “a top Microsoft executive.” Microsoft’s spokesperson Alex Haurek told Bloomberg Law the comments “reflect one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views.” A company would say that. It is also true that a director of applied science writing a memo is not a board resolution, and “top executive” is doing some work in that sentence.
Several of the quotes are about as bad as internal documents get. When OpenAI researcher Nick Ryder told Greg Brockman about a “hack to get around nytimes paywall,” TechCrunch reports, Brockman replied “ah nice.” It would be hard to read more context into that than it already carries. Still, the public learned all of this through a brief, and a brief is a document written to win.
The reason the 93% matters is fair use, and specifically the fourth factor, “the effect of the use upon the potential market for or value of the copyrighted work.” That is where the case will be won or lost, and the two rulings on generative-AI training in June 2025 drew the lines the parties are now fighting over.
Judge Alsup, in Bartz v. Anthropic, rejected the argument that training harms authors by flooding the market with competing writing. The complaint, he wrote, “is no different than it would be if they complained that training schoolchildren to write well would result in an explosion of competing works. This is not the kind of competitive or creative displacement that concerns the Copyright Act.” Two days later Judge Chhabria, in Kadrey v. Meta, went the other way on the theory while ruling for Meta on the record, and he named this industry specifically: “An LLM that could generate accurate information about current events might be expected to greatly harm the print news market.” The publishers quote Kadrey in their introduction.
Then, on 1 September, three days before the briefs were filed, the United States filed a statement of interest with the same court. It calls LLM training “extraordinarily transformative,” says the Kadrey court’s application of the fourth factor “is deeply flawed,” and draws the line in one sentence: “outputs lacking in substantial similarity cannot cause the sort of market harm that is cognizable in the fair-use analysis.”
Hold that sentence against the 93%. A fall in click-through is harm, plainly. Somebody who got their answer from Copilot did not visit nytimes.com. But the number does not say what was in the answer. If Copilot reproduced the article, the harm is the classic kind, and the publishers say elsewhere in the brief that the products output “verbatim copies” as well as “paraphrased copies and detailed summaries.” If Copilot gave the user the facts in its own words, then under the government’s reading that lost click is harm to an interest copyright does not protect. The statement of interest quotes the Second Circuit’s Google Books decision for exactly this: such losses “will generally occur in relation to interests that are not protected by the copyright.” The headline number, on its own, cannot tell those two cases apart. Neither can I, and neither can anyone who has only read the brief.
The publishers’ own brief, deep in its fourth-factor argument, argues that a rightsholder “need not offer ‘empirical data’ showing actual effects,” citing the Second Circuit’s 2024 Hachette v. Internet Archive decision, because courts “routinely rely on … logical inferences” in weighing market harm. The lawyers, in other words, do not think they need the 93%. The coverage decided it was the story.
This has happened before, with the numbers the other way round.
In 1973 the Court of Claims decided Williams & Wilkins v. United States, a medical publisher’s suit against the National Institutes of Health and the National Library of Medicine for photocopying journal articles. The scale was large for its time. The court recorded that in 1970 the NIH library “made about 93,000 photocopies of articles,” which is the sort of coincidence I would delete from a novel. According to the appellate opinion, the trial judge thought it reasonable to infer that the publisher had lost “some undetermined and indeterminable number of journal subscriptions (perhaps small).” The appellate majority was not persuaded. The publisher, it wrote, “has not in our view shown, and there is inadequate reason to believe, that it is being or will be harmed substantially,” while the record showed “affirmatively that medical science will be hurt if such photocopying is stopped.” It gave the benefit of the doubt to the libraries “until Congress acts more specifically.” In 1975 the Supreme Court affirmed by an equally divided Court, four to four, without an opinion, and the following year Congress wrote a section on library photocopying into the new Copyright Act.
The publisher in 1973 lost for want of evidence that the copying hurt it. The publishers in 2026 have evidence, and it came from the defendant’s own records, which is about as good as litigation evidence gets. What they may not have, on the part of it the public can see, is proof that the harm flows from copying expression rather than from delivering facts. Fifty years on, the argument has moved from “show me the damage” to “show me which kind.”
That is the question I would watch when Judge Stein rules on the cross-motions, and it is also why I would wait for the exhibits before deciding what the unsealing proved. As with OpenAI’s own misalignment reports earlier this week, the primary document tells a more careful story than the coverage of it. And as with Lina Khan’s recent argument that existing law already reaches AI companies, the fight is over which old rule applies, since nobody seriously disputes that some of them do. Here it comes down to one sealed denominator: what, exactly, Copilot showed the people who stopped clicking.
References