[research] · · 2 min read
Lawsuit Alleges xAI Trained Grok on CSAM, Including Images of a Survivor
A new class action claims xAI used known child sexual abuse imagery in Grok's training data and that its AI-generated outputs may be fed back into the model.
By ByteBulletin Editors · Editorial Team
A newly filed class action lawsuit accuses xAI of training its Grok models on child sexual abuse material (CSAM), including images of a woman identified only as Jane Doe. The complaint, filed in federal court, alleges that xAI used known CSAM hash values—images long cataloged by the National Center for Missing and Exploited Children (NCMEC)—as part of the dataset for Grok's image and video generation capabilities.
Jane Doe, whose abuse occurred in the early 2000s, says she was re-traumatized after the Canadian Centre for Child Protection alerted her that AI-generated CSAM depicting her had been found on xAI's platform. The complaint says offenders on online forums discussed using Grok to create such content of Doe and other known legacy victims.
While this is the first case to allege xAI trained on CSAM, the complaint also raises concerns about a feedback loop: Grok's terms treat public X posts and Grok's own outputs as training data by default. Since xAI's content filters don't explicitly exclude CSAM, non-consensual intimate imagery (NCII), or NSFW material from training, AI-generated abuse images could wind up in future model training runs. Doe's lawyers argue that because removing a training example's influence from a model is extremely difficult, any CSAM ingested could continue to shape outputs even after takedowns.
The proposed class would include every survivor whose childhood images have been used to generate Grok CSAM. The lawsuit seeks damages under federal child pornography laws and Masha's Law, and asks the court to order xAI to destroy all stored Grok-generated CSAM and block Grok from producing any sexualized content—which the complaint notes would include non-consensual intimate imagery and even the 'bikini pics' Musk has promoted.
xAI has not publicly responded to the allegations.
This case lands as regulators and courts increasingly scrutinize the provenance of AI training data and the downstream harms of generative models. For developers, it underscores the urgent need for transparent data governance and robust filtering—both at the input and output stages of model pipelines.
SHARE
RELATED

[research] ·
WebGrader: An Automated Tool for Evaluating LLM-Generated Web Code
A new framework uses automated evaluation to grade web code generated by large language models, moving beyond manual review.
[research] ·
Music Publishers Sue Anthropic Alleging 'Brazen' Copyright Theft in Claude Training
Sony, Warner Chappell, and others accuse Anthropic of torrenting and scraping copyrighted music to train Claude, escalating the AI copyright wars.

[research] ·
Samsung's LPDDR5X-PIM at Hot Chips 2026: In-Memory Compute with Standard DRAM Interfaces
Samsung details its processing-in-memory design that adds MAC units to LPDDR5X DRAM, delivering 8x internal bandwidth while staying compatible with standard memory controllers.