Overview
- Multiple outlets on Tuesday reported that contractors working on OpenAI projects were removed or fired after using large language models to complete reviewer tasks that were meant to be done by humans.
- Internal documents obtained by reporters explicitly forbid reviewers from using LLMs or named tools such as GPTZero and Grammarly to write feedback or comments and instruct reviewers to spot AI use by looking for patterns like repetition and unusually fast completion.
- Mercor, a vendor that supplies reviewers, confirmed its contracts ban LLM use and said it removes experts when misuse is confirmed, while OpenAI declined to comment on the specific firings.
- The reviewers handle real ChatGPT prompts and outputs at large scale, which raises privacy risks because automated redaction can miss personal details and an opt-out only blocks future, not past, data from training pipelines.
- There is no published evidence that these incidents have measurably harmed OpenAI models yet, but researchers warn that feeding AI-generated evaluations into training can over time degrade model quality in a process sometimes called model collapse.