OpenAI said they sicced this agent army on Navier-Stokes on Sept 1st, while only a couple of days earlier OpenAI's Noam Brown happened to reply to a tweet saying that they had already tried to solve all the Millennium Prize problems and failed... So, it seems either the previous attempt didn't have the training to succeed, or was just not given the compute to do so.
Once OpenAI heard that Navier-Stokes was solved, this caused them to immediately revisit the problem and throw a ton of compute at it, apparently using a more (very) recent model than what they had tried before. What we don't know is just how recent this model was, and therefore what it may have been trained on. Buckmaster/Levant had apparently been working towards this for at least a year, and made their "forced" blow-up breakthrough on August 15th.
Presumably any anonymized prompts that are being trained on are part of pre-training, so older, but once OpenAI had heard that Navier-Stokes had been solved and wanted to revisit it, it seems possible they may have done a few weeks of incremental RL training on anything Navier-Stokes adjacent they could come up with, in addition to then throwing unlimited compute at it, now confident that there was something to find.
>The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”
>The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”
what benefit do they get from making the statement? they could just say nothing. saying it and having it be untrue opens them to legal issues that are not worth the risk for this nothingburger.
And conceptually novel approaches to outstanding problems are the sort of thing that a retrain should pick up on, because they would be hard to compress into what it already knows.
Openai said that a new model became available to them during this. But that could mean anything from a big new base model to a LoRA, fine-tuned on a few dozen prompts...
It seems that in this case OpenAI are suggesting that the researchers whose work they scooped were using OpenAI models with an account setting that allowed OpenAI to train on anonymized prompts.
It seems that Buckmaster and Levant (who is an Anthropic employee) were rather naive in the amount of trust they had in OpenAI, with Buckmaster going so far as to contact OpenAI's Sébastien Bubeck to discuss what they were working on and clarify that contrary to rumor this was a private collab.
In reality you can't have things like infinite velocity, so if Navier-Stokes is predicting blowups, then isn't this a problem with Navier-Stokes not handling some edge cases, not a problem with reality that engineers need to be concerned about?
With 10,000 agents and $20M of compute this is just brute force search.
It's a bit like telling 10,000 kids there's an easter egg hidden over there, pointing to one corner of your yard (or having "heard a rumor" it was hidden in that corner).
If you have $20M to spend on your problem, then yes, AI brute force search is an option, but unless you know a solution is possible (as OpenAI did here), you may still be wasting your money.
You jest and that is OK. Brute force search is not something you can do over math problems of that difficulty or anything with combinatorial complexity.
To me it feels closer to taking the top 10k human mathematicians on a large retreat for a year and having them self organize to collectively solve this problem—not kids and easter eggs.
I'm not joking. Compare to a super-human MCTS system like AlphaGo or Stockfish - once you condense the expertise of your top 10K world experts into a board evaluation or policy function, then the rest is brute force.
Whether this type of agentic swarm approach can be considered closer to MCTS (search), or closer to a less structured GOFAI blackboard type approach (perhaps more like your mathematician retreat) I'm not sure - I don't think they've released any details of the prompt(s) and how these agents were collaborating and building on each others work.
The other part of my easter egg analogy is the direction to "look over there", corresponding to OpenAI specifically asking their hoard of mathematicians to work on Navier-Stokes since they knew it was solvable/determinable, and they certainly had the public work that Buckmaster/Levant were building on as further direction, as well as perhaps their prompts. Unlike Buckmaster/Levant, this wasn't just a couple of humans with a university research grant budget, this was apparently a not-so-small team at OpenAI (says Buckmaster, per a group call he had with OpenAI), with an unlimited budget, so it's hardly surprising (or in the least bit impressive) that they were able to duplicate and surpass their work.
You don't need to speculate here, since much of the story is not being disputed.
The recent ground-breaking work on Navier-Stokes "blow-ups" was done over a period of years by mathemtaticians Diego C´ordoba and Luis Martınez-Zoroa.
NYU professor Tristan Buckmaster and Anthropic employee (& mathematician) Levent Alpoge took the above work as a starting point, and over a year with LLM assistance developed a blow-up proof under certain conditions.
Buckmaster: "We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra. The latter was only used for writeups and auditing our arguments."
Buckmaster says he thinks that Martınez-Zoroa, whose work this all builds on, deserves the Fields Medal for his work.
OpenAI claim that on Sept 1st they heard a rumor the problem has been solved (which happened on August 15th), and then decided to re-solve it themselves using a 2-week old model, then later reached out to Prof. Buckmaster and Levant to come to some agreement to co-publish.
OpenAI: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ". In other words, not only did they deliberately choose to tackle a problem they heard had already been solved (in turns out only partially solved), but they may have done so using a model that was aware of the successful way to attack the problem.
It seems there are three potential scandals here:
1) OpenAI by their own admission chose to try to scoop mathematicians who they had heard had already completed a proof
2) OpenAI may have used a model that had seen "de-identified" messages indicating the direction to take
3) An OpenAI employee essentially threatened to "ruin the career" of the NYU professor who had been working on this if he did not cooperate with them
The direct plagiarism possibility, 2), while it should be a warning to anyone using OpenAI's models, doesn't need to be true for OpenAI to have benefited from the researcher's work. It's enough that they heard Navier-Stokes had been solved and could then go out with their swarm of 10,000 agents and $20M of compute to hunt out the latest research and brute force it.
Magnus Carlson once said that if he wanted to cheat all it would take would be for someone to indicate to him (a wink from someone in the audience perhaps) when a position warranted more time to be spent on it (because there was something important to be found if he did). It seems that, at absolute minimum, this is what OpenAI did here, although in context of math this is not cheating - the "wink" was a rumor, originating from who knows where, that a proof existed (but had not yet been published) and therefore there was potential to rush in and scoop rights to publish or co-publish.
The rumors were that Anthropic had solved Navier Stokes and was sitting in it to maximize IPO hype. It was all over twitter. In the first place the intent was never to scope mathematicians, but their closest rival.
The proof turned out to be different from what Levant and Tristan was doing.
An Astra sized model takes months to train - they are not saying 2 weeks to train from scratch. The only only interpretation of this "2 weeks" claim that is consistent with reality is that they mean 2 weeks of additional training on top of whatever their starting point was, so it's more like this:
|--- N months of base model training -->|--- X months of post-training -->(Astra?)|-- 2 weeks more training (on Navier-Stokes adjacent material, perhaps)--> this "new" model
If there is any truth to this timeline, then presumably it just means an additional 2 weeks of RL training on Astra.
"Twice as capable in mathematics" just means they found some problems that Astra couldn't solve, or make progress on (who knows how they chose to define "capable", or "twice as" for that matter), then put those 2 weeks of training in to focus on those gaps.
At this point, focused on their IPO, the best way to interpret OpenAI press releases is "what is the least this can mean, without being an actual lie". They are not shy - if there was a more impressive claim they could make, they would have made it.
>If there is any truth to this timeline, then presumably it just means an additional 2 weeks of RL training on Astra
OpenAI finished a larger pre-train (rumors are it's the largest since GPT 4.5) in late August (not Astra). Presumably, this is post training on top of that since it lines up.
>They are not shy - if there was a more impressive claim they could make, they would have made it.
I would think it would have the exact opposite effect.
Why would anyone use OpenAI models for anything commercially valuable, or where secrecy is important, when it appears that if OpenAI "becomes aware" that you are doing so they may try to compete with you?
Not only did OpenAI, by their own admission, rush to re-solve Navier-Stokes once they heard the rumor that it has been solved (the rumor being that it was Anthropic that had done it), but they are leaving the door open ("we cannot rule out that") as to whether the model they used to do it had been trained on the anonymized date from the researchers who's approach they ended up copying.
Terrance Tao has recently said as much for mathematics - that there appears to be a trend (not just this Navier-Stokes incident) of the AI companies going after math problems wherever there is an "rumor" of progress, and that he thinks this may sadly result in breakthrough mathematics being conducted in secret to avoid this.
The rush to steal another researcher's thunder is bad enough, but it also appears that one of OpenAI's employees acted in a very thuggish manner to try to threaten the professor who had been working on this not to publish and to co-operate with their telling of the story.
>but they are leaving the door open ("we cannot rule out that") as to whether the model they used to do it had been trained on the anonymized date from the researchers who's approach they ended up copying.
This is just lawyer speak. Maybe there was a small reward from a possible thumbs up on any of the chat sessions. Open AI have no way of knowing if that happened or not and the chances that, if this did happen, that it had anything to do with their solution of Navier Stokes is extremely unlikely.
>to try to threaten the professor who had been working on this not to publish and to co-operate with their telling of the story.
There was never a threat to not publish their own work with whatever credit to whoever. This was about the offer to be a lead author on the paper that Open AI authored, an invitation that was not extended to Levant.
Yes they know the timelines, which they've explained. No they don't know exactly what they train on. Even I don't, and my little experiments are nowhere near OpenAI scale. This is par the course for ML. Plus it would kind of defeat the purpose of de-anonimization if they could.
So OpenAI weren't even working on Navier-Stokes, but heard the rumor that "someone else" (Anthropic) has solved it, so rushed in to steal their thunder.
There are bound to be a bunch more results like this, in math, physics, chemistry, and now that we essentially have a DeepBlue for math, a DeepBlue for physics, etc, these results are going to come.
SOME of the problems that have eluded humans are going to turn out to be low hanging fruit that are susceptible to this type of brute force (10,000 agents on a supercomputer running for 7*24 hours straight) AI search.
I'd be more impressed if OpenAI found their own problems to solve, rather than rushing in to re-solve one once they heard it was already solved (and therefore not so hard).
Even under your interpretation, OAI pushed a button and solved NS. Yes, that is very impressive. Are you kidding me? Imagine building an automated system that can solve NS.
"I was shown a prompt and told the internal research model had simply been
given the problem statement. Levent had been told by Sebastien “very little
human input” had been used. This turned out not to be true. Over the course
of the call, as members of their team sent Sebastien corrections and details over
their internal chat, it emerged that an entire team had been working on the
problem".
The evidence we have (not much) leaves the facts severely under-constrained. All the below are consistent with what we observe:
- OpenAI lying their face off for marketing purposes
- Buckmaster sore about getting beaten to NS, misrepresenting what was told
- Buckmaster not being an expert in ML, not understanding what he was told and hearing what he wanted to hear as it vindicated him
- Others...
OAI claimed that none of the people on their team were domain experts. So it's more akin to Deep Blue than Stockfish (actually slightly better than Deep Blue which did hire domain experts) but nevertheless represents a highly sophisticated automatic theorem prover.
Even if true, I don't see why this is an issue. Are they not allowed to work on problems others are working on? Did Anthropic get first dibs on this problem? Competition is good. And I don't exactly have tons of sympathy when the other side is just a leading AI lab. It's not like it's some scholar who dedicated his life to this problem.
The "steal their thunder" is interpretation. What I'm saying is that you believe they solved NS on a lark to bully some other researchers, and that this is not impressive?
What's impressive for a human and for an AI are two different things.
Magnus Carlson had a peak ELO rating of almost 2900.
Would you be impressed with someone with an ELO of 3700?
Would you still be impressed if I told you it was Stockfish?
OpenAI didn't go looking for a tough-for-an-AI problem to solve - they went looking for one that looked like it was easy since it they had heard it had already been solved.
As a developer yes, especially given that it runs on a PC, and DeepBlue in it's day was really more impressive since is used custom ASICs.
But, I assume the Stockfish developers aren't comparing themselves to Magnus.
Let's see if OpenAI, or someone else, can get these sort of physics/math results out of a desktop PC - that would also be an impressive piece of engineering!
Possibly after being given the significant part of the solution from actual human researchers. Which they then bullied. And they beat them to the finish line only because they heard rumor and threw everything at the problem. It doesn't look good for openAI in any way. I see more reasons to avoid using them rather than use them from this story.
reply