The Copyright Endgame: Authors Deliver a Heavy Blow to OpenAI and Microsoft in Landmark AI Training Lawsuit

Executive Overview

The legal battleground defining the intersection of artificial intelligence and intellectual property has reached a critical boiling point. In a sweeping and aggressive legal maneuver, a consolidated group of aggrieved writers—including high-profile literary figures and class-action plaintiffs spearheaded by the Authors Guild—have filed a comprehensive motion for partial summary judgment in the U.S. District Court for the Southern District of New York. Presided over by Judge Sidney Stein, this pivotal proceeding targets tech titans OpenAI and Microsoft, accusing them of building the foundational infrastructure of their multibillion-dollar generative AI models upon a bedrock of mass digital piracy.

The plaintiffs’ motion cuts straight to the core of modern machine-learning ethics, demanding that the court rule unequivocally before any trial takes place: OpenAI copied their copyrighted works without authorization, and this widespread appropriation cannot under any legal standard qualify as "fair use." Covering a targeted selection of 194 copyrighted book titles, the filing seeks a definitive finding of liability rather than immediate monetary damages.

However, the implications of this legal salvo reach far beyond the immediate financial stakes for the individual authors involved. In the motion, the plaintiffs articulate an existential dread that resonates throughout the global literary and publishing communities: AI-generated books are flooding the marketplace, threatening to render human authorship economically unviable. By allegedly sourcing millions of illegal e-books from notorious shadow libraries via peer-to-peer torrent networks, scrubbing the digital fingerprints of their illicit origins, and deploying personnel who openly fantasize about using algorithms to finish the unfinished works of deceased or aging masters like George R.R. Martin, OpenAI and its primary financial backer, Microsoft, face an escalating reckoning.

As technology companies double down on their defense—arguing that transforming copyrighted text into statistical weights for neural networks is a legally protected transformative process—the courts are being forced to draw hard lines. With billions of dollars, the future of the generative AI industry, and the livelihoods of millions of creators hanging in the balance, this New York courtroom showdown promises to rewrite the rules of digital copyright for generations to come.


Detailed Chronology of the Consolidated New York Proceedings

To fully understand the gravity of the recent summary judgment motion, it is necessary to retrace the complex legal trajectory that brought these disparate lawsuits under a single judicial umbrella in Manhattan.

The Genesis: California Lawsuits and Cross-Country Migration

The seeds of the current litigation were sown in 2023, when several independent copyright infringement lawsuits were filed against OpenAI on both coasts of the United States. Among the earliest and most notable filings was the Tremblay and Silverman lawsuit, initiated in California by a group of authors who recognized that their copyrighted books had been scraped and ingested into OpenAI’s large language models without their consent or compensation.

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

Despite early efforts by OpenAI to dismiss the complaints, the Tremblay and Silverman action successfully survived a partial dismissal phase, establishing a crucial legal foothold that allowed the plaintiffs to press forward. Recognizing the overlap in defendants, legal theories, and discovery demands, subsequent actions—including a massive class-action lawsuit spearheaded by the Authors Guild, as well as a separate complaint filed by a group of prominent nonfiction writers who became the first to name Microsoft as a co-defendant—were bundled into a unified proceeding in the Southern District of New York under Judge Sidney Stein.

The Smoking Gun: From "LibGen" to "Books1" and "Books2"

At the heart of the plaintiffs’ new summary judgment filing is a meticulously documented trail of evidence alleging that OpenAI went to extraordinary lengths to obscure the illicit origins of its training data. According to the motion, when OpenAI constructed the early iterations of its revolutionary GPT models, it did not purchase legal copies of millions of books. Instead, the company allegedly turned to Library Genesis (commonly known as LibGen)—a notorious shadow library and piracy haven explicitly flagged by the Office of the United States Trade Representative (USTR) for widespread copyright infringement.

The authors claim that OpenAI employees downloaded mass collections of copyrighted books via torrent files from LibGen. Internal nomenclature within the company initially reflected these sources, with datasets explicitly cataloged as "Libgen1" and "Libgen 2." However, as the legal and public relations risks of large-scale scraping became apparent, OpenAI allegedly altered its documentation. In the academic and technical papers introducing GPT-3, the company quietly relabeled these datasets with the more nondescript, innocuous-sounding monikers of "Books1" and "Books2."

The paper trail of concealment did not stop at mere renaming. According to the court filing, OpenAI took the unprecedented step of entirely deleting its LibGen-derived training corpora in the summer of 2022 amid mounting legal anxieties. The plaintiffs emphasize in their brief that these represent the only two training datasets OpenAI has ever intentionally scrubbed and deleted from its archives—a move the authors characterize as a desperate attempt to destroy evidence of institutionalized, mass piracy.

The Gogineni Controversy: Replacing George R.R. Martin

Beyond the mechanics of data acquisition, the plaintiffs’ motion introduces damaging human-element evidence highlighting an institutional culture within OpenAI that views human writers not as partners or creators to be respected, but as obsolete obstacles to be automated away.

The motion highlights the actions and public statements of Tarun Gogineni, an engineer hired by OpenAI in 2022 to spearhead efforts to enhance the literary quality and generative writing capabilities of its models. In social media posts that have since drawn intense scrutiny, Gogineni openly acknowledged that his research mission would inevitably displace human authors. He casually dismissed this outcome as "acceptable economic disruption."

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

Most egregiously for the plaintiffs, Gogineni specifically targeted one of the lawsuit’s named plaintiffs: George R.R. Martin, the world-renowned author of the epic fantasy series A Song of Ice and Fire, which inspired the blockbuster HBO television adaptation Game of Thrones. In a 2025 social media post—made nearly two years after Martin formally joined the copyright lawsuit—Gogineni declared that his explicit research mission was to train GPT models to write the final two unreleased books of Martin’s legendary fantasy series. Adding insult to injury, Gogineni tweeted that even if Martin were to "die early," GPT-5 would successfully "autocomplete his series."

For the plaintiffs, these statements serve as a smoking gun, demonstrating that OpenAI’s business model was intentionally designed to cannibalize the exact creative works it misappropriated, replacing human creators with synthetic echoes generated by machine-learning models.


Supporting Context & Metrics: The Scale of the Infringement and Industry-Wide Fallout

The motion filed by the New York authors is part of an escalating, multi-front war being waged by the creative industries against Silicon Valley’s data-scraping practices. To appreciate the magnitude of the legal challenge, one must examine the broader ecosystem of litigation and the financial architectures supporting it.

The Microsoft Connection and Vicarious Liability

A critical dimension of the New York proceedings is the inclusion of Microsoft as a primary defendant. The authors argue that Microsoft cannot hide behind its status as an external investor, asserting that the tech giant bears direct vicarious liability for OpenAI’s copyright infringements.

Over the course of their partnership, Microsoft has funneled an estimated $13 billion into OpenAI across three major funding agreements signed in 2019, 2021, and 2023. The plaintiffs argue that this massive financial infusion was paired with deep operational integration. Microsoft provided the vast cloud-computing infrastructure (Azure) necessary to train OpenAI’s massive models, exercised supervisory control over OpenAI’s commercialization strategies, and stood to profit immensely from the deployment of technology built upon pirated literary works. By enabling and financially profiting from OpenAI’s data acquisition strategies, Microsoft is legally culpable for the downstream infringement, according to the brief.

The Expanding Front: Newspapers, Media Outlets, and Creative Guilds

The authors’ class-action suit is far from an isolated incident. Over the past several years, the legal dam has broken, with virtually every sector of the publishing and journalism industries taking legal action against generative AI developers.

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

In a synchronized display of industry alignment, plaintiffs including The New York Times, the Daily News, and the Center for Investigative Reporting have submitted their own combined summary judgment motions against OpenAI and Microsoft. These media organizations face parallel threats: their proprietary archives, investigative journalism, and daily reporting have been scraped to train chatbots that can synthesize and regurgitate news content without driving traffic or subscription revenue back to the original publishers.

Judicial Precedents: The Bartz v. Anthropic Caveat

The legal battleground over "fair use" is nuanced and evolving. AI companies frequently lean on landmark intellectual property rulings to argue that ingesting copyrighted text to train algorithmic models falls under the legal umbrella of transformative fair use.

However, recent judicial rulings have introduced vital caveats that undermine OpenAI’s blanket defenses. The plaintiffs’ brief heavily cites Bartz v. Anthropic—a pivotal 2025 California ruling which established that while the abstract concept of training an AI model on text might possess some fair use characteristics, the method of acquisition matters immensely. Specifically, that court ruled that downloading copyrighted books from illegal shadow libraries like LibGen when legal commercial avenues are readily available is "inherently, irredeemably infringing."

The authors argue that OpenAI’s decision to bypass legal purchasing channels in favor of peer-to-peer torrent networks strips the company of any credible fair use defense. Furthermore, they contend that scraping their specific books was never strictly necessary to create a generalized, foundational language model, rendering the unauthorized duplication legally indefensible.


Official Positions and Counter-Arguments

While OpenAI has yet to formally respond in court to the specifics of this latest summary judgment motion at the time of publication, the company’s overarching legal strategy is well-documented through its previous filings and cross-motions.

OpenAI’s Defense: Fair Use and Vanishing Regurgitation

In cross-motions for summary judgment filed simultaneously with the authors’ actions, OpenAI maintains a resolute defense. The company argues that its use of copyrighted books constitutes fair use as a matter of law. OpenAI’s legal team asserts that training a neural network to recognize statistical patterns, syntax, and semantic relationships in human language is a profoundly transformative process. Rather than substituting for the original books or acting as a direct market competitor, the AI model analyzes text to learn the mechanics of human communication.

OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

Furthermore, OpenAI contends that instances of generative models "regurgitating" substantial verbatim passages of copyrighted books are vanishingly rare anomalies rather than systemic features. The company maintains that its models are designed to generate original statistical predictions, not to act as digital copy machines dispensing pirated e-books to end-users.


Future Outlook: What Lies Ahead for AI and Copyright Law

The convergence of these high-stakes lawsuits in the U.S. District Court for the Southern District of New York represents a watershed moment for both the technology sector and the global creative economy.

If Judge Sidney Stein and other presiding judges rule in favor of the authors and media plaintiffs, the foundational business model of generative AI could face catastrophic disruption. Companies like OpenAI and Microsoft could be forced to scrub their training datasets entirely of unauthorized copyrighted material, delete models trained on illicit data, and pay astronomical statutory damages for past infringements. More importantly, it would establish a mandatory licensing framework, forcing tech firms to negotiate financial agreements with writers, publishers, and journalists before ingesting their lifeworks into neural networks.

Conversely, a sweeping victory for OpenAI on the grounds of fair use would cement a legal landscape where vast quantities of human creativity can be freely harvested by technology conglomerates without consent or compensation. For writers like George R.R. Martin and millions of other creators worldwide, such an outcome would validate their darkest fears: that the digital age has birthed an unstoppable technological apparatus designed to consume their labor, automate their artistry, and leave them economically obsolete.

As the litigation proceeds through the courts with millions of dollars and the future of human authorship hanging in the balance, one reality remains certain—the legal war over artificial intelligence training data is only just beginning.

Leave a Reply

Your email address will not be published. Required fields are marked *