NVIDIA Board Sued Over AI Trained on Pirated Books
AI Copyright · Lawsuit Filed
NVIDIA Shareholder Sues Jensen Huang and the Board Over AI Trained on Pirated Books, YouTube Videos and Cloned Voices
PublishedAugust 20, 2026
An NVIDIA shareholder has sued the company's own CEO, officers and directors, claiming they signed off on building NVIDIA's AI models from pirated books, scraped YouTube videos and voiceprints taken without consent. Because it is a derivative suit, any money would go back to NVIDIA — but the three class actions underneath it are the ones aimed at authors, video creators and voice professionals.
This article describes a stockholder derivative complaint. Everything below is an unproven
allegation drawn from that filing. NVIDIA and the individual defendants have not been found
liable, no class has been certified, and there is nothing for the public to claim. This page is
informational and is not legal advice.
What Is This About?
An NVIDIA stockholder filed a verified derivative complaint on July 31, 2026 in the U.S. District Court for the Northern District of Illinois, captioned Berliner v. Huang, No. 1:26-cv-09153. The 108-page filing names chief executive Jen-Hsun "Jensen" Huang, chief financial officer Colette Kress, three other executive vice presidents and the full slate of NVIDIA directors. NVIDIA itself is named only as a nominal defendant, which is standard in this kind of case.
The core allegation is that the defendants adopted and approved what the complaint calls an "ask forgiveness not approval" approach to AI training data — knowingly building NVIDIA's large language models, video models and voice models on material the company had not licensed, and then leaving NVIDIA exposed to the copyright and biometric-privacy lawsuits that followed. The shareholder brings claims for breach of fiduciary duty, along with federal securities claims under Sections 14(a), 10(b) and 20A of the Exchange Act, and demands a jury trial.
Nothing here is a consumer settlement. There is no fund, no claim form and no deadline. The reason it matters to readers is what sits underneath it: three separate proposed class actions against NVIDIA brought by book authors, YouTube creators and Illinois voice professionals, all of which the derivative complaint leans on as evidence that the board should have acted sooner.
StatusComplaint FiledDerivative action · no response from the defendants located as of August 20, 2026
Who Benefits If It SucceedsNVIDIA itselfRecovery in a derivative case goes to the company treasury, not to a consumer class
Can I Claim?No — nothing to claimNo class, no claim form and no deadline in this case or the three related class actions
What a Derivative Lawsuit Actually Is
A stockholder derivative suit is not a class action. The shareholder is not suing for her own losses; she is suing the executives and directors on behalf of the corporation, on the theory that the corporation was harmed by their conduct and that the same board cannot be trusted to sue itself. That is why NVIDIA appears on both sides of the caption — as the party allegedly injured, and as a nominal defendant.
It also explains why the complaint spends pages arguing that a pre-suit demand on the board would be futile. Under Delaware law a shareholder normally has to ask the board to bring the claims first. The complaint argues that step can be skipped because a majority of the directors face a substantial likelihood of liability themselves, sat on the board throughout the relevant period, and approved the disclosures at issue. Whether that argument holds is one of the first things a court will decide, and demand-futility rulings end a large share of derivative cases before the merits are ever reached.
The practical takeaway for a reader: even a total win here produces no payment to the public.
What the Complaint Says About NVIDIA's Training Data
The filing divides the alleged conduct across three families of NVIDIA models. All of the descriptions below are the shareholder's allegations.
Books and the NeMo Megatron language models
The complaint alleges NVIDIA's NeMo Megatron models were trained on "The Pile," a compiled dataset published by EleutherAI that includes a subcollection called Books3 — roughly 196,640 books the complaint says were sourced from a private piracy tracker. It further alleges NVIDIA used the SlimPajama dataset, itself derived from RedPajama, which also carried Books3 before that component was pulled in October 2023 over reported copyright infringement.
The most striking passage concerns Anna's Archive. Drawing on material described in the authors' case, the complaint alleges that in the fall of 2023, facing a deadline tied to its developer day and unable to strike licensing deals with publishers, NVIDIA approached the shadow library about including its collection in pre-training, was warned about the nature of that collection, and proceeded anyway with management's "green light" — ultimately obtaining access to roughly 500 terabytes of material. NVIDIA has not conceded that account, and no court has ruled on it.
YouTube video and the Cosmos world model
For Cosmos, NVIDIA's text-to-video foundation model, the complaint alleges the company drew on three research datasets built entirely from YouTube: HD-VG-130M (about 1,549,408 source videos), HD-VILA-100M (about 3,098,462) and HowTo100M (about 1,238,911). Those datasets, the filing says, distribute only identifiers and timestamps — so using them requires downloading each clip from YouTube directly.
That is where the Digital Millennium Copyright Act theory comes in. The complaint alleges NVIDIA used the downloader tool yt-dlp together with virtual machines that rotated IP addresses, and says internal Slack messages in a channel called "#cosmos-dataset-creation" show employees discussing YouTube's terms of service and IP blocking while continuing the work. It also alleges datasets licensed for academic or non-commercial research were used commercially anyway.
Voices and the commercial speech models
The third strand concerns NVIDIA's voice products — Magpie TTS Zeroshot and Magpie TTS Flow, FUGATTO, PersonaPlex, and the Canary and Parakeet speech models. The complaint alleges these were trained on hundreds of thousands of hours of human speech, from which the models extracted the acoustic signature that identifies an individual speaker, and that NVIDIA neither identified the source speakers, nor gave written notice of the purpose and duration of collection, nor obtained a written release — the three things Illinois' Biometric Information Privacy Act requires.
The filing makes a point that distinguishes voice claims from ordinary data claims: because several of these models were published as open-weight releases, the complaint argues the biometric data and the product have become the same thing, and that the product has already been distributed.
The Disclosure and Buyback Claims
Beyond the training data, the complaint alleges NVIDIA's annual reports, proxy statements and Code of Conduct told investors a story the defendants knew to be incomplete — that the company would "deliver trustworthy AI models that comply with privacy and data protection laws" and used only competitive data that was "publicly available or licensed to us."
The shareholder points to a change in wording as circumstantial support: that language appeared in NVIDIA's 2024 and 2025 proxy statements and, she alleges, was absent from the 2026 proxy. The complaint characterizes the removal as a tacit admission. That is the plaintiff's interpretation of an editing decision, not an established fact, and companies revise proxy language for many reasons.
The complaint also targets capital returns. It alleges the board authorized successive repurchase programs — $15 billion in May 2022, $25 billion in August 2023, $50 billion in August 2024 and $60 billion in August 2025 — and that NVIDIA bought back roughly $13.258 billion of its own stock between June 2024 and January 2026 at prices the shareholder says were inflated by the incomplete disclosures. A separate count under Section 20A alleges Huang sold NVIDIA shares during the same window while the company was repurchasing.
The Three Class Actions Underneath It
This is the part that matters to authors, creators and voice professionals. The derivative complaint is largely a repackaging of three proposed class actions already pending against NVIDIA:
Nazemian v. NVIDIA Corp., No. 4:24-cv-01454-JST (N.D. Cal.) — filed in March 2024 by novelists who say their books were in the training data. In May 2026 Judge Jon Tigar granted in part and denied in part NVIDIA's motion to dismiss, allowing direct and contributory copyright infringement claims to proceed. Surviving dismissal is not a finding of liability.
Ted Entertainment, Inc. v. NVIDIA Corp., No. 5:25-cv-10287-EJD (N.D. Cal.) — filed November 26, 2025 by YouTube creators, centered on the DMCA anti-circumvention theory. An amended complaint was filed in March 2026. The same company is behind a parallel case against Meta, which OCA covered in Ethan Klein's company suing Meta over AI video scraping.
Rogers v. NVIDIA Corp., No. 1:26-cv-05478 (N.D. Ill.) — filed May 12, 2026 by a group of Illinois broadcast journalists, podcasters, audiobook narrators and voice actors under BIPA, part of a coordinated set of filings that also named other large technology companies. The voiceprint theory is close to the one at issue in the Meta voiceprint BIPA class action.
The derivative complaint also references several additional copyright suits filed against NVIDIA during 2026 involving YouTube-hosted works, a Microsoft-created 3D dataset and a music catalog. None of the cases in this section has a certified class or a claim form.
NVIDIA is a Delaware corporation headquartered in Santa Clara, California, and the two copyright class actions were filed in the Northern District of California. This derivative case was filed in Chicago. The complaint grounds venue in an alleged NVIDIA office in Champaign, Illinois and in the BIPA claims involving Illinois residents' voiceprints.
Venue and forum are routinely contested in derivative litigation, and many companies have forum-selection provisions that push these claims elsewhere. Where this case is heard is a live question, and it may be resolved before anything about training data is.
What Happens Next
The defendants have not yet responded. In a derivative case the first substantive fight is usually over demand futility and the pleading standard, not the underlying conduct, and derivative suits are also frequently stayed while the related class actions move forward. As of publication, OCA had not located a public statement from NVIDIA or any individual defendant about this complaint.
The events more likely to produce news for readers are the ones in the three class actions: a class certification ruling, a summary judgment decision on fair use, or a settlement in the authors', creators' or voice cases. Those are the proceedings that could eventually create something a person could actually file for. We will update this page as the docket moves.
Read the Complaint
Frequently Asked Questions
Can I get money from this NVIDIA lawsuit?
No. This is a stockholder derivative action, which means the shareholder is suing on behalf of NVIDIA rather than for herself. If the case succeeds, any recovery goes to NVIDIA's own treasury, not to a class of consumers. There is no class, no claim form and no deadline for the public.
Has a court found that NVIDIA used pirated books to train its AI?
No. In the separate authors' case, Nazemian v. NVIDIA Corp., a federal judge in California granted in part and denied in part NVIDIA's motion to dismiss in May 2026, letting direct and contributory copyright claims move forward. Surviving a motion to dismiss means the allegations were adequately pleaded, not that they were proven. NVIDIA has not been found liable in any of these cases.
I am an author, a YouTube creator or a voice actor. Is there anything for me here?
Not in this derivative case. Three separate proposed class actions against NVIDIA cover books, YouTube videos and Illinois voiceprints. None of them has a certified class or a claim form, so there is nothing to file today. Anyone who thinks their work or voice was used should keep their own records and follow the dockets rather than sign up anywhere.
Why is a case about a California company filed in Illinois?
NVIDIA is a Delaware corporation headquartered in Santa Clara, California. The complaint asserts venue in the Northern District of Illinois based on an alleged NVIDIA office in Champaign, Illinois and on alleged Biometric Information Privacy Act violations involving the voiceprints of Illinois residents. Venue allegations are frequently contested, and the defendants have not yet responded.
What does the complaint say about NVIDIA's stock buybacks?
The complaint alleges the board authorized repurchase programs while the disclosed risks were incomplete, and that NVIDIA bought back roughly $13.258 billion of its own stock between June 2024 and January 2026 at prices the shareholder says were inflated. That is an allegation about harm to NVIDIA itself. It is not a claim on behalf of individual investors, and no court has evaluated it.