OpenAI is facing renewed scrutiny regarding the origins of the data used to train its mathematical capabilities, as a second researcher publicly challenged the company’s transparency. The allegations come just days after another mathematician accused the firm of benefiting from unpublished scholarly work.
Andreas Thom, a mathematician who raised his concerns on the social media platform Mastodon, stated that interactions he and his colleagues had with ChatGPT prior to OpenAI’s recent high-profile announcement may have contributed to the model’s success in mathematics. Thom characterized the company’s behavior as unethical and dishonest, specifically citing a lack of clarity surrounding the provenance of its training data.
The incident adds to an ongoing controversy over whether AI developers are appropriately crediting academic research during model training. Critics argue that the rapid advancement of these systems often relies on intellectual property shared in confidence or published in preprint form, raising questions about fair use and academic integrity in the age of artificial intelligence.
Honestly, I just wish they’d cite their sources properly. It’s not that much effort, OpenAI.
Mathematicians are finally pushing back. We need clear attribution rules before these models consume everything we write.
How can anyone verify this? OpenAI never releases their raw training data, so proof is nearly impossible to get.
This is deeply concerning. Using unpublished conversations without consent sets a dangerous precedent for all researchers.