OpenAI Accused of Training on Opted-Out Research Data
A mathematician says he explicitly opted out of OpenAI training data in June, was told by OpenAI that his data 'did not happen' to be used, and then found evidence that it clearly was. The thread is skeptical of the specific evidence presented, with several commenters pointing out the claim rests on 'someone somewhere had a discussion with AI about the topic' which is thin. But the underlying concern, that OpenAI's opt-out mechanisms are unreliable or dishonest, is getting serious engagement.
This lands on top of an already live debate about whether researchers and academics can trust AI labs with unpublished work. The Moonshot/Kimi story running in parallel today, where Moonshot was caught serving Claude responses instead of its own model and collecting the exchanges for training, adds another data point. The theme connecting them: AI companies are playing loose with data provenance and consent, and researchers are starting to notice systematically rather than anecdotally.
The pattern here is that trust is fracturing at the research layer. If academics start treating AI labs like adversaries with respect to unpublished work, that changes how labs get access to frontier knowledge and how AI-assisted research collaboration evolves.
So what?
If you're building products that ingest user-generated content or position yourself as a trustworthy AI infrastructure layer, the bar for demonstrable, auditable opt-out compliance just went up. 'We have a checkbox' is not enough. Researchers and enterprise customers are going to start demanding proof.