Overview
- The 404 Media exposé published Sept. 14–15 revealed Project Lily, an OpenAI initiative that hires hundreds of contracted reviewers to read real ChatGPT conversations and rate multiple model replies on a 1–7 scale.
- OpenAI says chats pass through an automated Privacy Filter that removes identifiers, but the company admits the filter can miss details and leaked dashboards show reviewers can see a 'user memories summary' that may reveal location, profession or past prompts.
- Users can turn off the 'Improve the model for everyone' toggle to stop future chats from being used for training, but the setting is on by default for Free, Plus, and Pro accounts and does not remove or change past data.
- Contractors are reportedly hired through third-party firms such as Crossing Hurdles and Mercor and are paid above $50 an hour to summarize user intent and flag style issues like 'AI-speak' or excessive flattering rather than perform full factual audits.
- The practice matches wider industry norms—companies including Anthropic and Google have similar human review—and the revelations are likely to renew regulatory scrutiny and calls for clearer user notice and stronger safeguards.