You can access the lawsuit from here.
A group of publishers and authors has filed a class action lawsuit against Google in a New York federal court, accusing the company of illegally copying millions of copyrighted books and journal articles to develop and train its Gemini artificial intelligence models.
Who filed the lawsuit: The plaintiffs include Hachette Book Group, Cengage Learning, Elsevier, author Scott Turow, and S.C.R.I.B.E. Inc. They allege that Google committed large-scale copyright infringement by copying copyrighted works from its own book services, downloading books and articles through web scraping, repeatedly reproducing them during AI training, and removing copyright management information (CMI) in violation of the Digital Millennium Copyright Act (DMCA).
The complaint opens with a sharp allegation against Google, saying, “Desperate to maintain its online dominance, Google abandoned its early motto of ‘Don’t be evil’ and engaged in one of the most prolific infringements of copyrighted materials in history.” It further alleges, “Google first illegally copied millions of books and journal articles… Google then copied those stolen works many times over to train its multi-billion-dollar generative AI system called Gemini.” According to the plaintiffs, “Google reproduced millions of copyrighted works without permission, without providing any compensation to authors or publishers.”
Use of Google’s own services: According to the lawsuit, Google copied books that had been provided for limited purposes through services such as Google Books, Google Play Books and Google Scholar, and later reused those works to train Gemini without obtaining fresh permission from publishers or authors. The complaint argues that these services were created to help users discover, search or buy books, not to supply training material for commercial AI systems.
The publishers also allege that Google copied books and journal articles from internet datasets created through web scraping, including material allegedly obtained from pirate websites such as Z-Library, OceanofPDF, WeLib and other sources, as well as content behind paywalls. They claim Google later copied these works multiple times while converting them into training datasets and model parameters for successive versions of Gemini.
Google allegedly knew the risks: The complaint argues that Google knew the legal risks. It cites internal documents that allegedly described using “Publisher Provided [] copyrighted books” from Google Play Books as “highly problematic for Google,” warning of “$10Bs-$100Bs in potential fines.” Another internal assessment allegedly stated, “Book publishers [are] likely to see LLM training on their books as copyright infringement. Could withdraw their content from Google Play Books [or] file a lawsuit against Google.” The lawsuit also quotes Gemini’s lead engineer as saying, “we don’t do deals for data we already have or already possess.”
The plaintiffs further claim Google’s internal documents show copyrighted books were considered necessary to improve AI performance. One document allegedly states, “It is important to emphasize that the team is keen to get access to high volume/high quality books asap and that format is critical.” Another reportedly concluded that “using only public domain books to train the model… resulted in lower performance.” The complaint also quotes an early Google AI developer as saying, “the current competitive landscape will force Google’s hand to develop AI faster… As a consequence, responsibility and ethical practice might be bypassed.”
Examples cited by the plaintiffs: To support its claims, the lawsuit cites several examples of Gemini’s outputs. It alleges the chatbot reproduced parts of economist N. Gregory Mankiw’s textbook Principles of Economics, generated a detailed summary of Scott Turow’s Innocent after being prompted by a user who said they did not want to buy the book, and produced an imitation of Lemony Snicket’s Who Could That Be at This Hour? that copied key creative elements. The complaint also cites Gemini’s response about N.K. Jemisin’s The Fifth Season, where the chatbot said, “Yes, the information included in that response comes directly from my internal training data.”
It also quotes Gemini’s “Thinking” mode stating, “My response is a definite ‘yes.’… the data includes the full text or sufficient extracts of The Fifth Season…” The lawsuit presents these chatbot responses as evidence, though Google is likely to dispute whether they accurately describe how the models were trained.
The publishers argue that Gemini now competes directly with the books used to train it. “The result is an AI system that competes directly with Plaintiffs’ and the Class’s works in the market,” the complaint says. It also claims, “Gemini can generate a 100-page murder mystery… in 20 minutes for a mere $0.39. No publisher or author can compete with that.” The lawsuit further alleges that Gemini sometimes encourages requests for copyrighted material by responding, “That’s a fantastic idea!” while suggesting prompts to generate additional infringing content.
Economic harm alleged: The plaintiffs allege Google’s actions have reduced book sales, undermined a growing market for licensing copyrighted works for AI training, and flooded the market with AI-generated substitutes for books and academic works. The complaint notes that Google already licenses some content for AI training from companies such as the Associated Press, Reddit and Shutterstock, but alleges it chose not to obtain similar licences from book publishers whose works it already had access to.
The lawsuit also links the alleged infringement to Google’s expanding AI business, noting that Alphabet reported its first $100 billion revenue quarter in October 2025 and said the Gemini app had crossed 650 million monthly active users. The plaintiffs argue Google has profited from AI built using copyrighted works without paying authors or publishers.
The plaintiffs are seeking certification of a class of affected copyright owners, damages for alleged copyright infringement and DMCA violations, disgorgement of Google’s profits, permanent injunctions to stop the alleged infringement, destruction of infringing copies where required, and a jury trial. The complaint concludes, “Copyright law applies to AI companies, including Google, with the same force as every other company that has complied with these laws for decades.”
Why it matters: The lawsuit adds to a growing legal battle over whether AI companies can use copyrighted works to train their models without permission. Two recent California rulings held that AI companies could claim “fair use” when training models on copyrighted works under US copyright law. However, a court ordered Anthropic to pay $1.5 billion for training its AI on pirated books. Because publishers filed the Google case in a New York federal court, a different judge will now decide whether Google’s fair use defence applies.
Unlike many other AI copyright cases, the publishers argue Google had access to their books only for limited services such as Google Books, which displays short snippets, and Google Play Books. They allege Google later used copies from those services to train Gemini without obtaining fresh permission, making the dispute distinct from cases involving content scraped solely from the open web.
Read more: