The most consequential AI copyright case in history just became the most consequential AI discovery dispute. The New York Times wants 8.1 million Copilot output logs in searchable format. Microsoft says producing them would cost millions and takes six weeks. The court has a deadline: August 21, 2026.
In December 2023, the New York Times filed the lawsuit that would define the legal relationship between AI companies and the publishers whose content trained their models. The case — New York Times v. Microsoft and OpenAI — alleged that the companies used millions of Times articles to train their AI systems without permission, reproducing content in ways that directly substituted for the original. It is the largest and most closely watched AI copyright case in history.
In 2026, it became something else as well: the most consequential AI discovery dispute in history. The News Plaintiffs filed a motion to compel Microsoft to produce approximately 8.1 million output logs containing the News Plaintiffs' names and web domains from Copilot — formerly Bing Chat — in a searchable format. The plaintiffs state that Microsoft is in a position to produce the logs but has delayed production as a litigation tactic.
The significance of the dispute extends far beyond the Times case. It is the first major court ruling on how AI output logs — the records of what AI systems generate in response to user prompts — must be produced in litigation. And the deadline for mediation reporting is August 21, 2026, with discovery timelines resuming immediately after. What happens in this case will define the discovery obligations of every AI company facing similar demands.
The Times and other News Plaintiffs are not simply asking for the logs — they are asking for them in searchable format. Microsoft has not disputed that the logs exist or that it can produce them. The dispute is about format, cost, and whether Microsoft's delay constitutes a litigation tactic designed to limit the plaintiffs' ability to use the logs as evidence.
The plaintiffs say the logs are directly relevant to proving two things: first, direct infringement — that Copilot reproduced News Plaintiffs' content verbatim or near-verbatim in response to user prompts; and second, "pink-slime" outputs — low-quality AI-generated text that mimics journalistic writing and directly competes with the News Plaintiffs' content in ways that cause measurable market harm. Without searchable logs, the plaintiffs argue, they cannot conduct the keyword searches and pattern analysis necessary to establish the scope of infringement.
The mediation deadline of August 21, 2026 — today — represents a potential pivot point. If mediation fails, discovery resumes on an accelerated schedule, and the motion to compel the 8.1 million logs will likely be resolved in the coming weeks.
The New York Times v. Microsoft dispute over 8.1 million Copilot output logs is not simply a fight between a newspaper and a technology company. It is the opening phase of a legal framework that will govern how AI output data is treated as electronically stored information in litigation — what it must contain, in what format it must be produced, and at whose expense.
Every organization that uses AI tools in its operations, that licenses AI-generated content, or that faces any form of AI-related litigation needs to understand what this dispute is establishing. AI output logs — the records of what a system generates in response to user queries — are ESI. They are subject to preservation obligations under Rule 37. They are producible under Rule 34. And the format in which they are produced — searchable or not — is a discovery dispute that courts are now being asked to resolve for the first time.
The Times case is the most visible example, but it is not the only one. As AI tools become ubiquitous in enterprise settings, the logs those tools generate — Copilot outputs, ChatGPT conversation histories, custom model interactions — are accumulating at a scale that makes their eDiscovery implications enormous. An organization that generates millions of AI interactions per month and has not considered those logs as a category of ESI is operating with a governance gap that litigation will eventually expose.
AI output logs are ESI. They are preserved, collected, and produced under the same rules as email and text messages. The format dispute in the Times case is simply the first time a court has been asked to define what that means at scale.
The New York Times is suing because Copilot allegedly reproduced its journalism. Microsoft is resisting searchable log production because 8.1 million logs in a structured format is an enormous discovery burden — and because searchable logs make it dramatically easier to prove the scope of alleged infringement. Both positions are rational from a litigation strategy perspective. Both illuminate something that every organization using AI tools needs to think through before litigation arrives.
The question is not complicated: if your organization uses AI tools that generate outputs based on user prompts, do you know what those logs contain, how long they are retained, and what you would be required to produce if those logs were subject to a discovery request? For many organizations, the answer is that they have never thought about it. AI tool logs are not managed with the same intentionality as email. They are not included in information governance policies. They are not part of legal hold checklists. And yet they may be among the most probative evidence in any litigation that touches on how AI was used in an organization's operations.
The Times case also illustrates the format dimension of AI log discovery that does not exist with most traditional ESI. Email can be exported in standard formats that are searchable by any litigation review platform. AI output logs may be stored in proprietary formats, distributed across cloud infrastructure, or structured in ways that require significant processing before they can be reviewed. The organization that understands its AI log architecture before receiving a discovery request is in a fundamentally stronger position than the one that discovers it for the first time when asked to produce eight million of them.
The New York Times v. Microsoft dispute is the most significant AI discovery battle currently before a federal court. But the questions it is raising — what AI output logs contain, how long they must be retained, in what format they must be produced, and who bears the cost of producing them — are questions that every organization using AI tools at scale will eventually need to answer.
At Sovereign Discovery, we are already incorporating AI output log governance into the information governance frameworks we build with clients. The tools are new. The legal obligations are not. If the data is created digitally, it is potentially discoverable. If it is potentially discoverable, it needs to be part of your information governance framework before a discovery request makes it urgent. The New York Times v. Microsoft case is simply making that urgency visible at an extraordinary scale.
The New York Times sued because it believed AI reproduced its journalism without permission. Microsoft resisted producing the logs that would prove or disprove that claim in searchable form. The court has a deadline of August 21, 2026. Whatever the outcome, the legal framework for AI output log discovery is being written right now — and every organization that uses AI tools will be bound by it.
AI output logs are ESI. They are not email. They are not documents. But they are discoverable — and every organization that generates them at scale needs to understand what that means before a request for 8.1 million of them arrives in their inbox.