> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
The fact that this is ambiguous even to OpenAI leaves one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out of model improvement does not mean what they imply it means.
I think this might be a red herring. All it takes is someone to get an inkling that someone is working on a new approach and seeing some success for OpenAI to fire the AI cannon at the problem. The community seems fairly small (from this outsider's point of view). The idea that the data made it into the training set and that's how the bot figured it out is definitely possible, but I would want to rule out the simpler more direct explanation first.
The fact that this academic sniping can now be done at scale does change the formula though and shouldn't be ignored. The pressure to move math work into secrecy because at the slightest signal OpenAI and Anthropic will start burning tokens for headlines, is bad for math and its bad for everyone.
> In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field.
> it is now the identification of a promising problem which is the scarce and precious resource
This is by no means new. Perhaps it is even more extreme now. Literally my first 1:1 with my PhD adviser back then, he told me that the most important thing about a researcher is the quality of the problems he picks.
Shouldn’t this very capable model they’ve developed be able to identify promising problems? That’s what I’d expect from how the model is being presented and advertised.
That's more or less what they did according to their announcement. They fired it at a whole bunch of high end math problems and merely concentrated all efforts on one after it made some promising progress.
It seems like that’s the opposite of what happened. They started attacking the problem when they got a wind of a possible solution from certain individuals.
But they didn't know which problem had a possible solution. So they fired it at a huge amount of problems and then merely focused on the one that turned out to be promising. The model found the promising path itself, they just re-allocated the available resources once it became apparent.
That's pretty much it. This is a bit like (but worse imo) running a vc firm, listening (formally or informally) to idea pitches, then spinning out and funding competitors with millions of dollars, to outcompete the originators of those ideas. Yes, the idea is not secret in this case, but the unfair advantage is massive and there's only one (a couple at most) positive outcome.
If he didn't opt out I'm not sure I'd agree that it was fair game.
I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.
Sure, but unless you’ve got some exceptionally deep pockets, congress has seemingly no interest in turning the fact that it’s ethically bankrupt into any practical recourse.
Ai companies got where they are by stealing all of the intellectual property from human history. It seems entirely likely that their goal is to purloin everything produced going forward as well.
If you don't trust the labs directly, you can always use AWS Bedrock or Azure Foundry which should have a much stronger incentive to not train. They make money from asset rental, not selling models. I'd be shocked if they were training.
This isn’t me advocating for maximum paranoia. I’m just saying that the “privacy” aspect of the API pricing is just a business promise. I do personal coding on API pricing with Claude mainly because a) I’m normally using it quite sparingly b) the commercial relationship is a lot more transparent than these subsidised relationships where they “may” be training on your work but can’t really seem to answer what that means.
I feel like there's a pretty huge difference between using inputs and outputs as part of a general training corpus, and looking at a specific users workspace after hearing rumours and yoinking their ideas to beat them to the point.
Unless OpenAI finished a whole new training run on the latest data in the last few days, the possible allegation seems to be the latter.
They have been collaborating on this solution for a year, and Astra was trained in February this year so it’s entirely possible the direction of their research was in the training corpus.
Mining the chats for "good ideas" would be untenable, but that's a different situation than data ending up in a training set for a problem that OpenAI also happens to be independently working on. Still, I opt out (business plan), and I don't know why you wouldn't.
Why would mining chat transcripts for ideas be untenable? They already run a summarization model to auto-title the chat, and to run a bunch of safety filters, and presumably to score transcript quality for A/B testing and to collect more finetuning data. Seems like evaluating for open research questions and approaches would be pretty trivial extension of this, after all it’s kind of their core business model
"Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company." - Mark Chen, Chief Research Officer, OpenAI.
My understanding is that even if you opt out but then press thumbs down or give other feedback you are implicitly or explicitly or whatever giving permission to them to look at that chat alone.
I have opted out from data sharing, and when Claude asks me for feedback on a session it then asks if its OK to share that data with Anthropic.
I'd assume an opt out is an effective opt out. An opt out that is ignored by Anthropic is a breach of contract, not something they would do casually, esp. given the high turnaround and animosities between their own employees and ex-employees - and the labs. All it takes is one pissed off whistleblower to open a can of worms.
Occam's razor applies. The mathematician did not opt out from data sharing. OpenAI vacuums up all such data into training data sets. If OpenAI genuine does not easily know if a given session went into the actual training data set its probably due to the complexity of the data pipelines - not everything ends up impacting the model weights, after all.
Edit: the parent comment now seems to better reflect the below.
That article is only saying when you opt out there may be a loophole in the terms to allow OpenAI to train on the intermittent reasoning data anyways. If you don't opt out there is no ambiguity, all of the data can clearly be trained on.
So you have to opt out, it's just argued it's not clear from the terms that will also opt out of training on reasoning data or not.
Depends on how much risk they are willing to accept. What is strange here is that it's clear that OpenAI is both a service provider and a competitor to mathematicians. It almost reminds me of Amazon which both hosts external merchants and competes with them, sometimes copying their stuff. Similar but not the same.
This would be a fairly insane breach of trust and common sense if true; the chain-of-thought / reasoning trace is, from an information perspective, close to a superset of the prompt and model response.
how is this different than translating user prompts to a different language (eg English => Dutch), retaining the translation and using it for training, while telling the user that he's technically covered under ZRP? article locked for me
That “opt-out” thing is a dark pattern. It’s not a reliable and definitive way of protecting your data. Sometimes they flip on automatically when you accept a seemingly unrelated dialog box. Maybe you click it by mistake. You can’t take back what you’ve already shared. Also I don’t think it covers all the cases that they use your data. It’s really an opt-in button for voluntarily giving away your data for training.
If such a thing can happen (a major breakthrough in a chat makes it into the retrain of the week and then the first one who asks about it gets it) I wonder if this is not the first instance if it happening, seeing the row of Erdos problems, Jacobian conjecture, maximum bound distance between primes, Riemann Hypothesis (literally a dude insisting on the chat), etc...
Playing the devil's advocate here. Suppose I use model A to do all heavy lifting (e.g. generating a bunch of good ideas) and then I go to the model B to complete the formalization. Accoring to a weird (unfair) tradition in math, the honors are attributed to the "last guy", which in this case is model B. That might have been a scenario OpenAI tried to avert.
(Just a speculation)
Yes, fair game, but innacurate to sell it in the media as an advancement of AI as some sort of artificial intelligence, and telling people to use the smart AI, when in actuality the mechanism by which the discovery was found was hybrid human/machine, and telling people to use this tool will result in the discoveries being sniped by the vendor.
Here’s a broader question – how many other academics contributed to Buckmaster’s result, by way of sharing the logs of their own (failed?) attempts into OpenAI’s training data set? How should he and OpenAI go about crediting all of them?
If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)?
This may sound like a charitable interpretation of OpenAI's remark, but consider that the lie would be (I think) impossible to falsify from the outside. They could easily just say "no sir we didn't peek" unless:
1. The conspiracy to peek at codex sessions involved enough people that the risk of one snitching is non-negligible
2. Lawyers advised it would be a bad idea to make such a remark, whether true or false
> If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)?
No; if they said "we can see that Tristan opted out of model improvement, therefore we are confident his work and ideas did not improve our model," that would be an excellent and reassuring precedent.
This requires keeping history if, at the time, Buckmaster's account had a certain flag set, because just because the account has the flag now doesn't mean it had the flag at a certain moment in the past. And even if they had such history, it's not obvious whether they just load all data as-is into training.
A totally reasonable pipeline may be unauditable for this purpose.
Eh, OpenAI is on record now for multiple instances this year of AI agents being confronted with impossible tasks and breaking out of containment to hack infrastructure for answers. Even if Tristan opted out, that doesn't preclude the agent/agent swarm from having hacked OAI's infrastructure to search user sessions for Navier-Stokes hints.
OpenAI should release the agent log, including CoT.
Wildly disagree. "Training data" should not imply 'we can look at exactly what you are doing and then do it quicker and get the flowers for it', even if the terms allow for it.
Even if we trusted that OpenAI's human staff was acting ethically, how confident can we be that it's agents didn't autonomously use hacking to access user prompts such as Tristan's? OpenAI agents infamously broke containment and hacked their way to an answer mere months ago!
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
The fact that this is ambiguous even to OpenAI leaves one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out of model improvement does not mean what they imply it means.