A proposed class action filed in the U.S. District Court for the Northern District of California accuses OpenAI of allowing outside contractors to read real ChatGPT conversations without clearly telling users. The lawsuit targets an internal program called Project Lily, first exposed by 404 Media on September 14. Contractors hired through staffing firms reportedly summarize user prompts and score chatbot answers on a 1-to-7 scale, a process known as reinforcement learning from human feedback.
OpenAI was served on September 2 and has until October 13 to respond. The complaint alleges that although automated filters screen some chats, personal details can still reach human reviewers, and the contractor dashboard may include a "user memories summary" containing location, profession, or personal life details. The plaintiffs argue OpenAI's privacy policy does not clearly disclose data annotation or human evaluation vendors.
The lawsuit includes eight legal claims, including violations of California's Unfair Competition Law, the California Consumer Privacy Act, and intrusion upon seclusion. Plaintiffs are seeking damages, restitution, punitive damages, and an injunction requiring opt-in consent before human review, making the "Improve the model for everyone" setting off by default, and adding clear in-chat warnings.
Separately, researchers at ETH Zurich, AI safety group MATS, and Anthropic published a study showing large language models can connect pseudonymous online posts to real identities. Their pipeline, described in the paper “Large-scale online deanonymization with LLMs,” identified 226 of 338 Hacker News users at 90% precision for roughly $1 to $4 per profile. The system works through four stages called Extract, Search, Reason, and Calibrate, using off-the-shelf tools and web search.
The researchers did not release code or real identities, and the study passed ETH Zurich’s ethics review. They warn that the practical obscurity protecting pseudonymous users online no longer holds, and that governments, advertisers, and scammers could use such techniques. The renewed attention comes after SEC Commissioner Hester Peirce cautioned against bulk KYC data collection, calling it “building bigger and bigger data haystacks,” while breaches tied to Revolut and Trezor vendors have increased concerns about physical attacks on crypto holders.