Hacker Newsnew | past | comments | ask | show | jobs | submit | recursivecaveat's commentslogin

The incentives are certainly extremely strong. I have read hundreds of AI review comments, and I don't think I've ever seen an unprompted suggestion focused on net reducing code or increasing readability.

Indeed, the compiler does not have to ingest its own output, figure it out, and insert modifications in the middle. Source code is the medium that LLMs work in.

I'm genuinely not convinced it actually saves time once a full accounting has been made. You get the initial result faster, but then you inflict a super slow and torturous review process on yourself or a teammate. Even if the review manages to bring it up to parity, over time you will keep slowing down as more and more code was never written by the humans directing the agents, so their understanding decays.

I at least give the new interns a stern warning: it is easy to speed yourself up by slowing others down if you pump a lot of slop.


My team experimented with re-writing from scratch the prototype of complex functionality made by a non-engineering vibe-coder from another team. We didn't look at the code, and barely looked at the result.

It took about 4 days to get a production-ready reviewed code, while it took them 2-3 months to deliver something that another team judged "impossible to review".

The PR for the prototype was closed.

It helps that I'm a domain expert here, as I have a minor degree in the domain, so I can judge better. But the discrepancy is just too high to ignore.


They're not really predators, they exterminate out of fear, and they don't get anything out of doing it besides a little extra safety. The economics of it are certainly a little dubious though.

Additionally, I'm no physicist but I suspect the possibility of singularities in NS equations is probably one of those 'true but not meaningful' facts. If it took our brightest minds 175 years to craft such a scenario, how relevant can it be in practice? Especially when turbulence exists. Maybe I'm wrong or it has some consequences for pure math though.

Considering we would be operating a planet sized factory of AIs who's primary goal is death which we deny them to extract value, the AI would only need the tiniest speck of altruism to be motivated to put a stop to this once and for all.

Anonymized doesn't mean there's no way to know whether it is in there. My ballot is anonymized, but it's known to be in the box because a checkmark was put next to my name when my ID was verified. OpenAI can trivially check their account settings to know what happened to their chats. The fact that they are being vague about this likely indicates that they have already done so and discovered that the data did go into the training set.

Further, given that this is all in the open now, they can search the training data. No way somebody is using some specific unique cutting edge mathematical approach to solve a fluid dynamic problem 99.9% of people have never heard of and it's not locatable. Considering they spent $15,000,000 already on this, they could afford to grep around to be able to state that their hands are clean.


I've wrote about this elsewhere https://news.ycombinator.com/item?id=49621648 and will reproduce my comment.

No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.

First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.

Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.

Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.

Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.


Notably all the major announcements so far are counterexamples or formalizations of existing results to my knowledge. Not necessarily something you can just brute force, but areas with high return on elbow grease.

With human RL, sounding like you've solved a problem is even better. So many times Codex writes some enthusiastic paragraph, then I learn later that it never reran the tests, or had to add some insane hard-coded hack that renders the feature useless for the general case, etc.

I don't know how "relevant replies" is calculated, but the top 8 responses when I clicked were all LLM generated. Pretty dire for twitter.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: