Here’s a broader question – how many other academics contributed to Buckmaster’s result, by way of sharing the logs of their own (failed?) attempts into OpenAI’s training data set? How should he and OpenAI go about crediting all of them?
There are probably thousands of serious academics taking a crack at millennium problems using AI every day. All those attempts are in the training data. And in fact the two researchers benefited from those attempts as well.
What else can they declare really? Yeah the model has training data from previous attempts. Alpöge and Buckmaster also similarly benefited from attempts before theirs.
I don't think OAI should be given the benefit of doubt. They are doing the research equivalent of front-running. Knowing where to look is one of the main challenges in research. Tristan's argument from his essay was that it is hard to brute force with a vanilla prompt (even for seasoned mathematicians) unless you knew very specifically what to mention i.e the search space would have been intractable even for OAI's compute budget.
"deidentified data" isn't much to go by. Say I prompted the internal model this way - "Hey there's a solution to a unsolved problem X. The solution uses a less known Method Y so don't bother wasting time with the usual methods. Take papers A, B and C as references. Oh btw, here's the last year's worth of data of all prompt sessions that mention this problem. Pay special attention to the ones that mention Method Y and sub-keywords Z,W".
This is obviously all speculation but the timing is very suspect. If OAI actually did this (and I suspect whatever they did is pretty much close to this), I think it is highly unethical.
> Tristan's argument from his essay was that it is hard to brute force with a vanilla prompt (even for seasoned mathematicians) unless you knew very specifically what to mention i.e the search space would have been intractable even for OAI's compute budget.
This is a bad argument. This is clearly not how it works. And unless Tristan is some truly alien-like savant (and maybe he is), what's necessary to initiate the AI's work already exists in countless published research papers and not exclusively in his head or notes. AIs can survey the sum total of all prior work on a problem and discern reasonable paths for inquiry.
Tristan is acting as if he's working off of an outdated model of AI, similar to primitive chess-playing models that winnowed the search space much more deterministically. If someone this intelligent truly doesn't get that this is not at all what AI is anymore, then maybe there's no hope that we ever understand it.
But I think he does realize this and he's flailing about for counterarguments from a place of bitterness and dejection, accepting even those that are too weak to be defensible. And that is very human and even forgivable.
"Given how seriously this would violate the most fundamental of academic standards, as well as taint the claimed capability behind this result, we take this issue very seriously, and we're launching a probe into identifying whether any of their research artifacts have entered our training set. We have further begun making changes to our UI/UX on all our surfaces, so that it is always clear whether any particular chat, or other user artifact, is eligible for being trained on."
They could have thought about the problem for like 2 minutes and not done this! I think that literally any academic mathematician could have explained to them, had they asked, why it is considered extraordinarily rude to react to rumors of research progress by desperately rushing to get there first.
> I think that literally any academic mathematician could have explained to them, had they asked, why it is considered extraordinarily rude to react to rumors of research progress by desperately rushing to get there first.
Pretty much all of math and science history is basically this pattern again and again. I'm sure all of that was rude as well.
Being scooped is not a new phenomenon, but the scooper's story is almost always that they were working on the problem independently or had some independent insight into it. By OpenAI's own account, they were inspired to start working on this by rumors that there might be Millennium Prize solutions to scoop.
Apologies if this is against the rules, but could I ask if you have some background in scientific research (maybe you could elaborate lightly on topics you've worked on)?
From my perspective, this practice is quite bad mannered, unusual, and heavily frowned upon, but I recognize it's possible that these stories might be more common in other fields. Still, I'd appreciate a strong sign that you aren't making these statements up based on secondhand accounts of what 'academia is usually like'.
CS PhD, used to be a professor, have worked at several industry research labs since.
> this practice is quite bad mannered, unusual, and heavily frowned upon,
Yes, it is.
Maybe you missed my point?
The fact that is "quite bad mannered, unusual, and heavily frowned upon" does not stop it from happening, and it is common for all high profile inventions and discoveries.
This does not disqualify the person doing the scooping, history remembers them as having the credit, and quietly forgets the person who was scooped.
If it's all business as usual and being scooped is no big deal, why was OpenAI in such a rush? They didn't have to launch this effort on the very day they heard the rumor, run "on the order of 10,000 concurrent agents", or try to coordinate announcement scheduling with Buckmaster in the middle of a long weekend. It seems to me that they understood very well this was not a "business as usual" announcement, and devoted huge amounts of money and focus to maximize the chance that they were first.
I think you're making the opposite conclusion than what I intended?
Being scooped is a big deal for the one getting scooped.
It has never been a big deal for the one doing the scooping. History is full of math and science results being scooped. For example, we keep calling it Pythagoras' theorem a few thousand years later.
> Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
It seems like this is going to be a PR nightmare, because they are now competing with their own customers. If you're using an LLM to help with your bright idea to cure cancer, you're going to have second thoughts about relying on OpenAI.
In OpenAI's case, if they were genuinely unsure, they wouldn't have said anything. "We cannot rule out" means they absolutely 100% for-sure did look at the existing prompts and bootstrapped from that, and they are trying to get ahead of the disclosure with this weasel-wording.
Also possible: we're 99.999% sure, but a lawyer said to be safe and strictly accurate, we should stick in a sentence in saying we can't be perfectly sure, since it's infeasible for us to prove it.
I promise you that if we took their work from ChatGPT and stuck in a bunch of weasel words to give the opposite impression while remaining technically true, I would quit on the spot.
Bundling in codex isn’t going to increase the download count on their website. Unless you’re assuming that every codex install downloads a fresh non-cached copy.
If I publish something, and disclose that I used AI for assistance, do I have to credit everyone who previously used the same AI to try the same problem? Because their prompts inevitably made it to the training data for my prompts?
reply