The motion
Microsoft filed for summary judgment on 4 September in the consolidated AI copyright litigation before Judge Sidney Stein in Manhattan federal court, arguing that training large language models on books and news articles is fair use and should be resolved as a matter of law rather than at trial. The consolidated case sweeps in claims by The New York Times, the Authors Guild and, among others, some 400 local newspapers.
The argument rests on discovery data rather than on doctrine alone.
What the logs show
Microsoft gave an expert retained by the news publishers 8.2 million Copilot conversation logs. The sample was not random: it was selected using keywords tied to the plaintiff outlets’ websites, which is to say it was filtered to be the portion of traffic most likely to contain their journalism.
Of those 8.2 million records, 59,545 contained at least 16 words in common with news content the model was trained on — under one per cent of the sample. On books, Microsoft’s filing puts the number lower still: 24 instances of reproduced passages across the same 8.2 million conversations, or 0.00029 per cent, with no matching output at all for 202 of the 212 books at issue.

Why the framing is the fight
The number is not really in dispute; what it means is. Microsoft’s position is that a rate this low, in a sample deliberately loaded against it, shows Copilot is not a substitute for the underlying journalism, and that training is transformative in the sense fair use requires.
Publishers have made a different argument throughout this litigation: that the harm is not primarily verbatim regurgitation but the use of the work to build a product that answers the questions readers would otherwise bring to the publication. On that theory, a low reproduction rate is beside the point. Nothing in Microsoft’s brief resolves that disagreement — it addresses the copying question the publishers also raised, and asks the court to treat it as the case.
The “16 words in common” threshold is worth noting too. It is a mechanical overlap test, and it counts matches that would not be infringing on any reading, which cuts both ways: it inflates the raw count while making the resulting figure a ceiling rather than a floor.

Where this sits
The filing lands two days after the US Department of Justice told the same court that training AI on news is fair use, a position that put the federal government on the side of the defendants in a case brought by a newspaper. Microsoft, meanwhile, began selling OpenAI’s GPT-6 Astra through its Foundry platform on 2 September, at $10 per million input tokens.
No US appellate court has yet answered the fair use question for model training, and district judges have reached conflicting conclusions. A ruling from Judge Stein on this motion would not settle it nationally, but it would be the most consequential decision so far in the newspaper cases — and if he denies it, the question goes to a jury.