NEW YORK / RankWire.AI / – Hachette Book Group, Cengage Learning and Elsevier have initiated legal action against Google concerning its Gemini artificial intelligence platform. Author Scott Turow and his company, S.C.R.I.B.E., have joined the class action lawsuit. The complaint was filed on July 10 in the U.S. District Court for the Southern District of New York. The plaintiffs accuse Google of copying millions of copyrighted books and journal articles without authorization during the development and training of Gemini. As of July 15, the court had yet to rule on the allegations or certify the proposed class.

According to the complaint, Google acquired material through Google Books, Google Play Books, and Google Scholar. Publishers and authors had provided works for specific purposes, including search functions, sales, and research access. The plaintiffs argue that these agreements did not permit broader commercial AI training. Furthermore, they allege Google downloaded extensive web-scraped datasets containing copyrighted works, some of which originated from known pirate sources and services protected by paywalls.
The 57-page complaint presents four claims under federal law. Three relate to alleged reproduction through Google services, web scraping, and the development or training of Gemini. The fourth claim invokes the Digital Millennium Copyright Act, asserting that Google removed or altered copyright management information from training materials. The filing also references internal discussions about using publisher-supplied books. One evaluation estimated potential fines ranging from $10 billion to $100 billion. These allegations have not yet been tested in court.
Proposed class includes registered works
The class being proposed encompasses owners of registered U.S. copyrights in qualifying books and journal articles. To qualify, books must have an International Standard Book Number (ISBN), and articles must have a Digital Object Identifier or International Standard Serial Number. The class covers works allegedly copied from Google services or obtained through web scraping. It also includes works purportedly reproduced during the development or training of Gemini.
Eligibility for membership also depends on registration timing. One criterion requires registration within five years of publication and prior to Google’s alleged reproduction or distribution. Another requires registration within three months of publication. The lawsuit excludes government entities, Google affiliates, certain court participants, and individuals who properly opt out of the class. The court must approve the class designation before the case proceeds to represent the broader group.
Complaint requests damages and an accounting
The plaintiffs seek statutory damages or actual damages for proven infringements and demand Google profits attributable to any confirmed copyright violations. They also request an injunction, legal fees, and a jury trial. The complaint does not specify a total damages amount but asks Google to disclose Gemini training data, collection methods, and model capabilities through a court-ordered accounting.
This accounting would identify the copyrighted works used in training Gemini and detail how Google collected, copied, processed, and encoded these materials. The plaintiffs additionally seek court-supervised destruction of unauthorized copies under Google’s control. Earlier, Hachette and Cengage sought to join separate Google AI-related litigation in California. The current case expands to include Elsevier, Turow, and S.C.R.I.B.E., focusing on claims related to Google services, web scraping, and Gemini’s training process.