By Legal & Technology Desk Published: October 2025 Main Facts Apple Inc. is facing a high-stakes class-action lawsuit filed in the U.S. District Court for the Northern District of California, accusing the tech giant of widespread copyright infringement in the development of its artificial intelligence systems. The lawsuit, spearheaded by neuroscientists, authors, and State University of New York (SUNY) professors Susana Martinez-Conde and Stephen Macknik, alleges that Apple utilized unauthorized, pirated books—most notably from a notorious "shadow library" dataset known as Books3—to train its proprietary AI architecture, including the Apple Intelligence platform, OpenELM (Open Efficient Language Models), and its foundational large language models. According to the legal filing, Apple bypassed standard licensing protocols and neglected creator compensation, allegedly masking its sourcing practices behind vague industry jargon such as "publicly available" and "open-sourced." The plaintiffs argue that while Apple aggressively secured multi-million-dollar content partnerships with major corporate entities like Shutterstock, Condé Nast, and NBC News, it treated individual authors and creators as fair game for uncompensated data scraping. The plaintiffs are seeking substantial statutory damages for willful infringement, a permanent injunction against the continued use of the contested models, and an unprecedented court order compelling the complete destruction of any AI models and training datasets built using their unauthorized copyrighted works. Chronology of Events To understand how the legal battle unfolded, a clear timeline connects Apple’s technical disclosures, dataset distributions, and the mounting legal pressure across the tech industry: October 2023: The controversial Books3 dataset—containing approximately 196,640 pirated books, including Martinez-Conde and Macknik’s international bestseller Sleights of Mind: What the Neuroscience of Magic Reveals About Our Everyday Deceptions—was removed from the hosting platform Hugging Face following widespread accusations of copyright infringement and "defunct" accessibility. June 2024: Apple officially acknowledged that its long-running web-crawling program, Applebot (which had operated in the shadows for nearly a decade), was actively harvesting internet data to train its foundational AI models. The disclosure arrived too late for creators to utilize opt-out protocols, as Apple had already finished its primary training phases. July 2024: Apple published its Foundation Language Model (FLM) research paper. The document outlined the company’s heavy reliance on licensed publisher data, curated open-source datasets, and web-crawled content. Crucially, the filing noted that licensed data was restricted to a secondary phase called "continued pre-training," leaving the core pre-training phase exposed to unverified, potentially infringing material. September 2025: The shadow of Bartz v. Anthropic loomed large over the legal landscape when a landmark preliminary settlement established the largest publicly reported copyright recovery in history, setting a militant precedent for AI developers accused of pirating textbooks and literature. October 9, 2025: Martinez-Conde and Macknik formally filed their class-action complaint against Apple Inc. in the U.S. District Court for the Northern District of California, citing systemic intellectual property theft and deceptive concealment of training data sources. Supporting Data & Technical Allegations The core of the legal complaint relies heavily on technical documentation published by Apple itself, alongside metadata recovered from open-source repositories and AI research papers. 1. The Shadow Library Connection: Books3 and The Pile The plaintiffs highlight that Apple’s own model cards and GitHub repositories for OpenELM explicitly state that its pre-training data relied on The Pile and a subset of RedPajama. The Pile, a massive dataset curated by EleutherAI, famously integrated Books3—a pirated library originating from the private BitTorrent tracker Bibliotik. Despite the removal of Books3 from Hugging Face in late 2023 due to legal liabilities, Apple’s models had already absorbed its contents. Shawn Presser, the creator of the dataset, previously admitted that developers were acutely aware of the severe copyright vulnerabilities shadowing the project from its inception. 2. The Applebot Scrape and Quality Filters Apple’s FLM paper detailed how its proprietary web-crawler, Applebot, gathered vast swaths of internet text. To ensure high performance, Apple deployed advanced "model-based classifiers" to filter out low-grade information. However, the lawsuit asserts that these very classifiers were trained on unlicensed, copyrighted material, compounding the cycle of infringement. Furthermore, the lawsuit introduces a novel allegation: Apple allegedly trained its AI systems on unauthorized digital copies of eBooks sold to consumers directly through the Apple Books ecosystem. The plaintiffs argue that the digital rights management (DRM) and purchase agreements governing Apple Books do not grant the corporation the right to repurpose purchased literature into internal AI training fuel. The Market for AI Data & Selective Licensing The economic disparity highlighted in the lawsuit underscores a growing rift between big tech and independent creators. The Booming Data Economy: Market analysts estimate that the commercial trade of licensed AI training data is currently a $2.5 billion industry, projected to skyrocket toward $30 billion within the decade. Selective Corporate Partnerships: While refusing to compensate independent scholars and novelists, Apple demonstrated that it recognizes the legal necessity of licensing by securing high-profile commercial deals. Most notably, Apple forged a multi-million-dollar agreement with Shutterstock (valued between $25 million and $50 million) for visual data, alongside ongoing financial negotiations with media powerhouses like Condé Nast and NBC News. The Stock Market Windfall: The complaint notes the stark contrast between Apple’s frugal data-sourcing methods and its financial gains. On the day Apple officially unveiled its Apple Intelligence ecosystem, the corporation’s market capitalization surged by more than $200 billion—marking the single most lucrative trading day in the company’s history. Broader Implications for the Tech and Publishing Industries The lawsuit filed by Martinez-Conde and Macknik is more than an individual dispute; it serves as a crucial test case for the future of generative artificial intelligence, copyright law, and the preservation of human-authored literature. 1. The Threat of AI-Generated Market Dilution Beyond the immediate issue of uncompensated theft, the plaintiffs argue that Apple’s deployment of generative AI directly undermines the economic viability of human creators. As AI models ingest copyrighted books, they generate summaries, derivative works, and competing texts that flood online marketplaces. This proliferation of low-quality, automated "sham books" threatens to crowd out original academic and creative literature, diminishing the long-term incentive for human authorship. 2. Legal Precedents and the Shadow of Bartz v. Anthropic The timing of this lawsuit is heavily influenced by shifting judicial attitudes toward AI copyright defenses. The legal team leans heavily on the watershed ruling in Bartz v. Anthropic, where a federal court drew a hard line against corporate data scraping, famously asserting that copying a textbook from a pirate website constitutes copyright infringement—full stop. With courts increasingly skeptical of broad "fair use" protections claimed by AI developers who ingest pirated literature, Apple faces an uphill battle. If successful, this class action could force major tech companies to completely scrub their foundational models, fundamentally altering how AI systems are built, verified, and monetized globally. Disclaimer: The articles, analyses, and legal summaries published on this platform do not constitute formal legal advice nor create an attorney-client relationship. The views expressed herein reflect ongoing legal proceedings and should be independently verified through official court docket filings (Case No. 5:25-cv-08695, N.D. Cal.). Post navigation The Retroactive Royalty Revolution: How Songwriting Credit Became a Negotiable Currency The Jurisprudence of Reality: Why Courts Must Rethink Executive Remedies in an Era of Institutional Defiance