Executive Overview
The legal battleground between human creators and artificial intelligence giants has reached a critical inflection point. In a sweeping and aggressive legal maneuver, a consolidated group of prominent authors—spearheaded by the Authors Guild and high-profile class actions—has filed a comprehensive motion for partial summary judgment in the U.S. District Court for the Southern District of New York. Presided over by Judge Sidney Stein, this multi-faceted lawsuit targets both OpenAI and its primary financial and computational backer, Microsoft, accusing them of copyright infringement on a monumental scale.
At the core of the plaintiffs’ recent filing is a striking assertion: OpenAI did not merely experiment with publicly available data; it built the foundational architecture of its GPT models on a foundation of industrial-scale digital piracy. According to the court documents, OpenAI bypassed standard licensing mechanisms and commercial avenues entirely to source training materials. Instead, the company allegedly relied heavily on torrented files downloaded from notorious illicit repositories, most notably Library Genesis (LibGen).
Furthermore, the motion moves beyond historical data acquisition to challenge the underlying economic philosophy of modern generative AI. Citing internal employee communications, the authors claim that OpenAI’s leadership viewed human literary displacement not as an unfortunate side effect, but as an intentional and acceptable form of economic disruption. By highlighting explicit statements from researchers targeting specific authors—such as A Song of Ice and Fire writer George R.R. Martin—the plaintiffs argue that the ultimate commercial goal of these models is to directly supplant the very creators whose works were appropriated to train them.
As Judge Stein weighs these arguments, parallel proceedings involving major journalistic institutions, including The New York Times, are converging on similar legal questions. With billions of dollars in market valuation, foundational intellectual property rights, and the future trajectory of machine learning innovation hanging in the balance, this case promises to reshape the legal boundaries of AI training data for decades to come.
Detailed Chronology of the Litigation
To fully understand the weight of the current summary judgment motion, it is necessary to trace the legal journey that brought these disparate claims before Judge Sidney Stein in Manhattan. Over the past three years, the intersection of copyright law and generative AI has evolved from a theoretical academic debate into a relentless wave of federal litigation.
The Genesis: 2023 Class Actions and Cross-Country Moves
The friction between authors and AI developers materialized into formal legal challenges in early 2023. Among the earliest filings was the Tremblay and Silverman lawsuit, initiated in California. This action brought together writers who discovered that their copyrighted texts had been ingested into OpenAI’s massive training corpora without consent or compensation.
Despite initial skepticism from various legal observers regarding how traditional copyright statutes would apply to transformative machine learning technologies, the Tremblay and Silverman action successfully survived a partial dismissal motion in early 2024. Shortly thereafter, strategic legal maneuvering led to the consolidation of multiple New York-based actions into a single overarching proceeding under Judge Sidney Stein.

This consolidation brought together:
- The Authors Guild Class Action: Representing a broad coalition of published writers seeking class-wide redress for widespread text ingestion.
- The Nonfiction Writers’ Suit: A distinct group of authors notable for being the first to explicitly name Microsoft as a co-defendant alongside OpenAI.
- The Transferred Tremblay and Silverman Action: Moving from its West Coast origins to join the New York docket.
The 2025 Motion for Summary Judgment
The litigation shifted from a discovery-heavy phase to a direct clash over liability this week, when the plaintiffs filed their motion for partial summary judgment. Eschewing a request for immediate monetary damages at this stage, the motion asks Judge Stein to issue a definitive pretrial ruling on two foundational questions:
- Did OpenAI copy the plaintiffs’ protected works without authorization?
- Can this unauthorized ingestion and reproduction legally qualify as "fair use"?
Covering 194 specific book titles, the brief paints a picture of a corporation aware of its legal exposure, systematically scrubbing its records while continuing to market models trained on disputed datasets. In response, OpenAI filed a cross-motion for summary judgment of its own, doubling down on its assertion that training neural networks on public text falls squarely within the boundaries of fair use as a matter of law.
The Allegations: Built on Mass Piracy and Concealment
The most explosive revelations within the plaintiffs’ recent filing center on the provenance of the data used to train OpenAI’s early language models, specifically GPT-3. While artificial intelligence companies have historically defended their data collection practices under broad banners of "public internet scraping," the authors allege a far more explicit reliance on underground piracy networks.
The LibGen Connection
According to the redacted court documents, OpenAI bypassed commercial databases, library acquisitions, and direct licensing agreements during the critical development phase of its early models. Instead, the company allegedly turned to torrent networks to download entire digital libraries from Library Genesis (LibGen), a platform long recognized by the Office of the United States Trade Representative (USTR) as a notorious global piracy haven.
The motion bluntly states: "OpenAI did not even buy the books it used. Instead, it began by torrenting books from the notorious and illegal pirate library Library Genesis, also known as LibGen."
For the plaintiffs, this distinction is legally fatal to OpenAI’s fair use defense. While some recent judicial rulings—such as the 2025 Bartz v. Anthropic decision in California—have carved out nuanced pathways suggesting that the subsequent processing of text for machine learning may lean toward fair use, those same rulings explicitly draw a hard line at the method of acquisition. Sourcing raw materials from a dedicated pirate repository to avoid licensing costs, the argument goes, is inherently and irredeemably infringing.

Systematic Concealment and Data Scrubbing
The court filings do not merely accuse OpenAI of sourcing pirated materials; they allege a calculated effort to obscure the origin of those datasets from the public eye and regulatory scrutiny.
The plaintiffs point to OpenAI’s early academic papers introducing the GPT-3 architecture. In initial internal and developmental documentation, researchers explicitly cataloged training corpuses under unambiguous titles: "Libgen1" and "Libgen 2." However, when preparing public-facing research papers and documentation, OpenAI allegedly relabeled these datasets into more nondescript, sanitized categories: "Books1" and "Books2."
Furthermore, the timeline of data retention forms a pillar of the authors’ case. The motion highlights that in the summer of 2022—as legal scrutiny surrounding generative AI began to mount—OpenAI made the deliberate decision to delete its LibGen-derived training files. The plaintiffs emphasize that these remain the only two primary training corpuses OpenAI has ever officially deleted from its historical architecture. To the legal team representing the authors, this sudden purge is a tacit admission of guilt, confirming that OpenAI’s internal stakeholders fully understood the illicit nature of the material they had utilized to bootstrap their commercial products.
Replacing the Creator: Economic Displacement and Internal Rhetoric
Beyond the technical arguments surrounding data ingestion, the summary judgment motion introduces a deeply personal human element: the stated intent of OpenAI personnel regarding the future of human authorship. The plaintiffs argue that OpenAI’s ultimate business model relies not just on learning from human literature, but on actively replacing the human creators altogether.
Tarun Gogineni and the "Research Mission"
To substantiate claims of economic displacement, the legal team highlights public statements made by Tarun Gogineni, an OpenAI researcher hired in 2022 to optimize and direct the writing capabilities of the company’s models.
According to the filings, Gogineni openly acknowledged that the models he was training would inevitably displace human authors. Rather than viewing this disruption as an unintended negative externality, internal communications and public statements allegedly framed it as an acceptable economic reality.
The controversy is brought sharply into focus through a specific target of Gogineni’s research: acclaimed fantasy author George R.R. Martin, who is a co-plaintiff in the ongoing litigation. In statements highlighted by the court, Gogineni articulated a vision where advanced iterations of OpenAI’s technology could bypass human creators entirely. Specifically, Gogineni noted on social media that his research mission included enabling GPT models to write the long-awaited final two volumes of Martin’s epic fantasy series, A Song of Ice and Fire (the literary basis for HBO’s Game of Thrones).

In a post that has drawn intense scrutiny from legal analysts and literary communities alike, Gogineni suggested that even if Martin were to pass away before completing his life’s work, subsequent models like "GPT-5" could simply autocomplete the series.
The Existential Threat to Publishing
The inclusion of these statements serves a vital strategic purpose in the authors’ legal brief. By demonstrating that key personnel view the technology as a direct substitute for specific, named human artists whose works were ingested into the training data, the plaintiffs dismantle OpenAI’s claims of harmless, transformative educational use.
"OpenAI’s GPT models pose an existential threat to those who write and publish books," the legal brief warns the court, pointing to real-world market conditions where AI-generated texts of varying quality are already flooding digital marketplaces, diluting discoverability, and undercutting the livelihoods of professional writers.
Supporting Context, Metrics, and Broader Legal Ecosystem
The legal pressure on OpenAI and Microsoft extends far beyond the confines of the Authors Guild class action. The broader landscape of intellectual property litigation involving generative AI is expanding rapidly, creating a multi-front war for Silicon Valley’s leading developers.
The Microsoft Connection: Vicarious Liability
A critical component of the current summary judgment motion is its direct targeting of Microsoft. For years, major tech conglomerates have sought to insulate their primary investment arms from the direct operational liabilities of their AI partners. However, the plaintiffs are pressing for Microsoft to be held vicariously liable for OpenAI’s alleged copyright infringement.
The legal basis for this claim rests on financial scale and operational oversight. Throughout successive investment agreements signed in 2019, 2021, and 2023, Microsoft has poured roughly $13 billion into OpenAI. The authors argue that this unprecedented level of financial integration gave Microsoft the legal right and practical ability to supervise OpenAI’s conduct, while directly benefiting from the commercial exploitation of the allegedly infringing models deployed across Azure and consumer-facing applications.
The Expanding Judicial Front
The New York proceedings overseen by Judge Stein are part of a coordinated wave of judicial challenges. Just days prior to the authors’ summary judgment filing, another massive coalition of plaintiffs—including The New York Times, the Daily News, and the Center for Investigative Reporting—filed a combined summary judgment motion against OpenAI and Microsoft regarding the unauthorized ingestion of journalistic archives.

Key Legal Metrics at a Glance:
- Primary Defendants: OpenAI, LLC and Microsoft Corporation.
- Lead Judicial Overseer: Judge Sidney Stein, U.S. District Court for the Southern District of New York.
- Scale of Current Filing: Covers 194 specific copyrighted book titles in the partial summary judgment motion.
- Financial Exposure: Billions of dollars in foundational investments (including Microsoft’s $13 billion commitment) and potential statutory damages.
- Parallel Actions: Ongoing litigation involving major national newspapers, media syndicates, and nonfiction writer associations.
Future Outlook: What Lies Ahead
As both sets of litigants await formal rulings from Judge Sidney Stein on their cross-motions for summary judgment, the stakes for the global technology and publishing sectors could not be higher.
If the court accepts the authors’ arguments that downloading from pirate repositories like LibGen strips away any potential fair use protections, AI developers will be forced to radically overhaul their data collection pipelines. Such a ruling would render retroactive mass-scraping legally indefensible, potentially requiring technology companies to purge existing models trained on disputed datasets or face catastrophic financial penalties for willful copyright infringement.
Conversely, should OpenAI successfully defend its practices under the banner of transformative fair use, the balance of power in the creative industries will shift decisively toward technology developers. Authors, journalists, and visual artists would find themselves with significantly diminished legal leverage to protect their intellectual property from being utilized as foundational fuel for artificial intelligence.
With millions of dollars in damages, corporate valuations, and the fundamental definition of human creativity hanging in the balance, this legal battle in New York is far from reaching its final chapter. Both sides are preparing for a protracted, high-stakes war of attrition that will ultimately establish the foundational rules governing artificial intelligence for generations to come.
