The High-Stakes Legal Battle Over AI Training: Authors File Landmark Summary Judgment Motion Against OpenAI and Microsoft

Executive Overview

The legal battlefield surrounding generative artificial intelligence has entered a decisive new phase. Over the past three years, creators, publishers, and copyright holders have watched nervously as foundational AI models were built and scaled on vast oceans of human-generated text. Now, a coalition of authors—including prominent names bundled into a high-profile New York federal proceeding—has escalated the conflict significantly.

In a newly filed motion for summary judgment overseen by U.S. District Judge Sidney Stein in the Southern District of New York, plaintiffs are asking the court to rule before trial that OpenAI and its primary backer, Microsoft, copied their copyrighted works without authorization, and that this unauthorized utilization fundamentally fails to qualify as legal fair use. The filing covers 194 distinct titles, seeking an foundational finding of liability rather than immediate monetary damages.

The legal arguments cut to the core of the generative AI economy. The plaintiffs do not merely argue that training models on copyrighted books is inherently infringing; they allege that OpenAI built the very foundations of its multibillion-dollar enterprise on intentional, mass digital piracy. According to the court documents, OpenAI bypassed standard licensing channels, downloading unauthorized copies of books via torrents from the notorious underground repository Library Genesis (LibGen). Furthermore, the motion alleges a coordinated internal effort to conceal these actions by scrubbing references to the source material from technical documentation and ultimately deleting the offending training datasets when legal exposure became apparent.

As tech giants and creative industries clash over the future of intellectual property, this motion—alongside parallel actions from major journalistic outlets like The New York Times—represents an existential threat to the current paradigm of AI development. With billions of dollars in investments, commercial valuations, and the fundamental integrity of human authorship hanging in the balance, the legal fight between creators and code promises to reshape the technological landscape for decades to come.


Detailed Chronology of the Litigation

To understand the gravity of the current summary judgment motion, it is essential to trace the legal trajectory that brought these diverse claims into Judge Sidney Stein’s New York courtroom.

The California Roots and Migration

The seeds of the current New York proceedings were originally planted on the West Coast. In 2023, author plaintiffs filed landmark class-action lawsuits in California, most notably the Tremblay and Silverman lawsuit. These initial complaints exposed how large language models (LLMs) ingested vast quantities of text without compensating or seeking permission from the creators.

While OpenAI successfully weathered early motions to dismiss certain aspects of the litigation, parts of the Tremblay and Silverman action survived judicial scrutiny. Rather than remaining fragmented across different jurisdictions, several of these foundational cases were eventually consolidated into a single, massive proceeding in the U.S. District Court for the Southern District of New York under Judge Sidney Stein.

This consolidation brought together multiple factions of the literary world:

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court
  • The Authors Guild Class Action: Representing a broad cross-section of registered writers whose works were allegedly digested by OpenAI’s algorithms.
  • The Nonfiction Writers Group: A specialized cohort of authors who made legal history by explicitly naming Microsoft as a co-defendant alongside OpenAI, tying the tech giant’s massive financial investments directly to the alleged infringement.
  • The Migrated California Cases: Including the surviving claims from the Tremblay and Silverman pipeline.

The Motion for Summary Judgment

This week’s filing marks a pivotal turning point. Moving for summary judgment means the authors are asking Judge Stein to bypass a traditional jury trial on specific legal questions, arguing that the undisputed facts demonstrate that OpenAI infringed upon their copyrights as a matter of law.

The motion focuses heavily on 194 specific titles, demanding a declaration of liability. The plaintiffs argue that the evidence unmasked during discovery leaves no factual ambiguity: OpenAI took their books without permission, used them to train foundational models, and cannot hide behind the legal shield of fair use.

Concurrently, OpenAI filed its own cross-motion for summary judgment, doubling down on its assertion that training general-purpose AI models on publicly available text constitutes transformative fair use and that the actual "regurgitation" of protected text by its models is statistically vanishingly rare.


Supporting Context & Evidence: Mass Piracy and the "LibGen" Connection

The most explosive revelations in the authors’ recent filing center on the mechanics of how OpenAI allegedly sourced its training data during the formative stages of the GPT architecture.

The Reliance on Library Genesis (LibGen)

According to the redacted summary judgment brief, OpenAI did not acquire its early training corpora through legitimate commercial purchases, library databases, or voluntary publisher partnerships. Instead, the company allegedly turned to the digital underworld.

The motion alleges that OpenAI:

"…did not even buy the books it used. Instead, it began by torrenting [REDACTED] books from the notorious and illegal pirate library Library Genesis, also known as LibGen."

At the time of these actions, Library Genesis was already a primary target for international intellectual property enforcement, prominently featured on the U.S. Trade Representative’s annual list of notorious piracy markets. The plaintiffs assert that OpenAI leadership and engineers were fully cognizant of the legal and ethical controversies surrounding LibGen when they integrated its holdings into their training pipelines.

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

Systematic Concealment and Data Scrubbing

The filing alleges that OpenAI was not merely negligent, but actively engaged in obfuscation to hide its reliance on pirated materials from the public eye and regulatory authorities.

Key elements of this alleged cover-up include:

  • Nondenominational Renaming: In the foundational academic research paper introducing GPT-3, OpenAI reportedly altered internal dataset nomenclature. Compilations previously designated as "Libgen1" and "Libgen 2" in internal documentation were sanitized and renamed to the far more innocuous and nondescript labels "Books1" and "Books2."
  • Targeted Deletion: As legal scrutiny mounted and awareness of potential liability spread within the company, OpenAI allegedly took drastic steps. In the summer of 2022, the company quietly deleted its LibGen-derived files. The authors’ motion emphasizes that these remain the only two training corpuses OpenAI has ever intentionally deleted from its archives, signaling a tacit recognition of wrongdoing by corporate insiders.

As the authors summarize in their brief, OpenAI essentially built the foundational architecture of its commercial empire on a bedrock of mass digital piracy.


Official Statements, Internal Rhetoric, and Economic Displacement

Beyond the mechanics of data acquisition, the summary judgment motion explores the underlying economic philosophy driving AI development: the explicit intention to automate and replace human creative labor.

The Tarun Gogineni Controversy

To demonstrate that OpenAI’s models were intentionally engineered to substitute human authors, the plaintiffs highlight public statements and internal goals associated with key personnel. In 2022, OpenAI hired Tarun Gogineni to spearhead efforts aimed at dramatically improving the writing quality and narrative capabilities of its models.

According to the legal filing, Gogineni was acutely aware that the models he was training would inevitably displace human authors, a development he reportedly characterized as an "acceptable economic disruption."

Most strikingly, Gogineni directly targeted one of the named plaintiffs in the lawsuit: George R.R. Martin, the celebrated author of the epic fantasy series A Song of Ice and Fire (the source material for HBO’s Game of Thrones). In 2025—nearly two years after Martin formally joined the copyright infringement lawsuits against OpenAI—Gogineni posted on social media detailing his professional "research mission." His stated goal was to train GPT models to write the long-awaited final two books of Martin’s masterwork.

Gogineni reportedly added that even if Martin were to "die early," upcoming iterations like "GPT-5 will autocomplete his series." For the plaintiffs, statements of this nature serve as smoking-gun evidence that AI companies view copyright law not as a boundary to be respected, but as an inconvenience to be bypassed on the road to total creative displacement.

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

The Threat to the Literary Ecosystem

The legal brief warns of an impending cultural and economic catastrophe for human writers. It notes bluntly:

“OpenAI’s GPT models pose an existential threat to those who write and publish books… AI-generated books of all types are already flooding the market.”

By training models on stolen texts and subsequently empowering those same models to generate derivative works at scale, AI companies are accused of short-circuiting the traditional economic engine that sustains human literature.


Legal Analysis: Fair Use, Contributory Liability, and Broader Industry Fallout

The legal battlefield in New York does not exist in a vacuum; it is part of a complex matrix of emerging jurisprudence across multiple U.S. circuit courts.

The Fair Use Debate and the Bartz v. Anthropic Precedent

OpenAI and its industry peers have consistently anchored their defense in the doctrine of fair use, arguing that transforming copyrighted books into numerical weights and vector embeddings within a neural network constitutes a transformative, non-infringing use under U.S. copyright law.

However, recent judicial rulings have introduced nuanced limitations to this defense. The authors heavily rely on the 2025 ruling in Bartz v. Anthropic out of California. While that court acknowledged that training AI models on text can theoretically qualify as fair use under certain circumstances, it drew a bright, unyielding line regarding how that data is acquired. The Bartz court ruled that downloading source materials from illicit pirate repositories when legal commercial alternatives exist is "inherently, irredeemably infringing."

The New York plaintiffs argue that because OpenAI sourced its data via LibGen torrents rather than authorized channels, its fair use defense collapses entirely. Furthermore, they assert that copying entire literary works was never strictly necessary to train a generalized foundational model.

Microsoft’s Vicarious Liability

Crucially, the authors’ motion expands the net of accountability beyond OpenAI to encompass its principal financial and technological benefactor: Microsoft.

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

The plaintiffs argue that Microsoft should be held vicariously liable for OpenAI’s copyright infringement. Under federal copyright law, vicarious liability attaches when a defendant has the legal right and ability to supervise the infringing conduct while maintaining a direct financial interest in the enterprise.

To prove this financial nexus, the authors point to Microsoft’s staggering investments in OpenAI, totaling approximately $13 billion structured across three major partnership agreements signed in 2019, 2021, and 2023. By providing the massive Azure cloud computing infrastructure, capital, and commercial integration necessary to scale OpenAI’s models, Microsoft allegedly profited directly from the unauthorized ingestion of copyrighted literature.


Future Outlook: A Defining Moment for the Information Age

The summary judgment motion filed in the Southern District of New York is far more than a localized dispute between aggrieved novelists and a Silicon Valley startup; it is a stress test for the legal frameworks governing intellectual property in the twenty-first century.

This action is unfolding alongside a coordinated onslaught from the traditional media ecosystem. Over the past year, major news organizations—including The New York Times, the Daily News, and the Center for Investigative Reporting—have filed consolidated summary judgment motions of their own against OpenAI and Microsoft, echoing similar grievances regarding unauthorized data harvesting and the generation of competing synthetic content.

As Judge Sidney Stein weighs the opposing motions for summary judgment, the stakes could not be higher. If the court rules in favor of the authors, it could invalidate the core training datasets of major generative AI models, forcing the tech industry into a retroactive licensing regime that would cost tens of billions of dollars and fundamentally restructure how foundational models are built. Conversely, a victory for OpenAI would cement a broad interpretation of fair use, signaling that technological progress and computational efficiency supersede traditional copyright protections in the age of artificial intelligence.

With countless millions of dollars, corporate valuations, and the cultural future of human authorship hanging in the balance, this legal war is destined to be fought tooth and nail through every level of the federal judiciary. For now, the literary world watches and waits, knowing that the outcome of Judge Stein’s eventual ruling will write the opening chapters of the AI era.

Leave a Reply

Your email address will not be published. Required fields are marked *