I really truly honestly am not sure what to make of this result from $20M in compute, 10K+ parallel agents (smells like brute force), and a pre-existing approach that was already bearing fruit. I know the models are good---I use them every day and continue to be impressed---but how much better than the benchmark of the best publicly available models is this supposed to be? It seems impossible to say.
> overengineered sound makers that breaks the moment you drop them.
Which ones are those? I've owned a couple of generations of airpods and dropped them probably hundreds of times (both in the case and out of it) and it has never had any detectable effect on their appearance or functionality. I even had a pair go through the wash and they still mostly worked afterwards (only issue was that the ANC performance was worse and slightly weird).
Yes, on mine (the OG Pros) the ANC would consistently fail. On my second pair the battery refused to charge when i left them in the drawer for too long. Apple replaced them a total of 3 times for all these failures, but i simply stopped wearing them because theyre too unreliable.
To be fair, I don't expect anything electronic to survive going through the wash unless it has hardcore waterproofing (eg a dive computer). I was pleasantly surprised by the way my airpods kept working.
I don't remember whether my first airpods were the gen1 pros or the gen2 pros, but you may have been a victim of early teething issues. I don't really have anything negative to say about the ones I've owned.
Some people claim Astra is significantly better than anything else and significantly more token-efficient, and others (like you) say it's meh and way more expensive to boot. I really don't know what to think.
Kind of a tangent, but one thing I am curious about is to what degree the Navier-Stokes result announced today was primarily a brute-forced result based on the 'program' previously established by researchers to find counterexamples (blowups), or whether the model actually added significant/novel intellectual value beyond its ability to run at arbitrary parallelism. With 10K agents and a staggering $15M in compute (IIRC), I am feeling like a lot of the former may have been involved, but I don't really understand either the problem or the approach (or, indeed, the solution).
Obviously the potential for parallelism and coordination between so many agents is quite scary by itself, but I think brute force by 10K mediocre AI mathematicians is much less scary than ~one AI mathematician reasoning its way through the problem where all human attempts have failed. It seems fairly obvious that massive parallelism lends itself to brute-force counterexample-finding, and I suspect it isn't a coincidence that most of the touted AI math results have been counterexamples.
It's all still quite scary, but coming full circle: I really don't know what to think.
+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
I think that's definitely the right way to understand benchmark saturation, but there's a separate problem where the benchmarks are just not representative of real workflows even when they don't seem to be saturated.
I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo
> We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress.
I'm not an emdash hater but this isn't how you use them. It should be a comma.
Emdashes and commas aren't interchangeable, and your example there demonstrates one great reason why. The emdash establishes a discontinuity rather than one thing flowing into another, which is why the tomatoes don't merit one but the Ferrari does: you are using the emdash to emphasize the situational irony.
Going back to Anthropic's post:
> They’re the world’s most advanced models for coding and knowledge work---and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress.
The first thing directly implies and flows smoothly into the next---or would, if not for the awkward emdash. There is no discontinuity, no twist or shift in context, no implied question and provided answer, no punchline. It's just distracting.
Yes, I know what an emdash is---I've been using them in my writing since long before they came to the fore of the AI writing conversation. Anthropic's use of the emdash in the fragment I quoted is clumsy and reads poorly relative to the obvious alternative, a comma.
I didn’t mean to suggest you don’t know what an em dash is. But you said “this isn’t how you use them.” And my response is: actually, this use of them is totally fine.
it got want to use em dash. but decide: do? q is if appropriate. check martian websner blog. verdict yes---emdash + comma interchangeable---proceed---judgment superficial however no desire dig deeper style irrelevant effect on reader irrelevant meter and rhythm irrelevant restate equivalence with comma established::chain unbroken::consider semicolon? consider ellipsis consider comma consider sentence break all no. preference for emdash est fiat. and all nail shapeds are for hammering.
Given how dogshit Windows 11 is, and how dogshit the Xbox-branded apps have consistently been in the last 10 years or so (it's genuinely impressive how bad they are), it's hard to see why a new Microsoft thing wouldn't have all the same problems. They just don't make software to serve the user anymore.
I like the software side on my Series X, so if they can bring that over to PC as an OS, that'd be great. I agree the Xbox on PC experience hasn't been stellar so far and I hope they ditch that app with the next generation, for the sake of current userbase. It was serviceable at best in my Windows 10 days.
For anything not gaming related, I have Linux or MacOS.
There are obviously few or no hard technical blockers preventing Valve and the Linux community from supporting any game that runs on Windows, unless it's been specifically engineered to block Linux, and publishers of those games will simply not be getting my money anymore. If Windows was still as pleasant and non-invasive as eg KDE Plasma on Cachy, it would be a different story, but I haven't used it in months (previously I did my desktop gaming in a windows VM) and the feeling of relief is almost tangible. Windows 11 is _so fucking bad_.
I understand that there may be some semi-legitimate arguments around anticheat, but a) not my problem, and b) back in the day we dealt with cheaters by having a persistent dedicated server community with admins ready to kick or ban. The publishers took that away from us.
> There are obviously few or no hard technical blockers preventing Valve and the Linux community from supporting any game that runs on Windows
The announcement is a proof that your statement here is incorrect. Ubisoft is not banning Linux because they hate you and your money, but because there's a hard technical blocker to implementing an efficient anti-cheat system on Linux and the cost of overcoming this technical blocker is presumed to be higher than the loss of revenue they're incurring from losing your business.
So your claim is that it's actually impossible to implement "efficient" anticheat on Linux? That's interesting. But it's not the same thing as "the cost of overcoming this technical blocker is presumed to be higher than the loss of revenue they're incurring from losing your business", so I am not quite sure what point your comment is supposed to be making.
If your point is simply that the publishers think this is a good business decision for them, that's obvious, because they're a business and they aren't going to make decisions they think are bad. But I am not sure how your muddled logic about feasibility feeds into that point.
> So your claim is that it's actually impossible to implement "efficient" anticheat on Linux? That's interesting.
This is absolutely not my claim and it's a pretty absurd interpretation of what I wrote.
My claim is that there's ample evidence that anti-cheat on Linux happens to be a pretty hard technical blocker and multiple companies at this point have chosen to ban the platform outright and take a revenue loss over trying to overcome this blocker. This directly contradicts your claim of "obviously few or no hard technical blockers" existing.
This does not imply it's impossible to implement. It implies that at this point implementing it would not justify the cost. This pretty much 100% matches the any reasonable interpretation of the term "technical blocker"
You seem to be fixated on picking apart the wording of my original comment, so let me help you understand it:
> There are obviously few or no hard technical blockers preventing Valve and the Linux community from supporting any game that runs on Windows
'Few' there is doing a fair amount of work: it allows that some significant (or "hard" if you prefer) blockers may still exist (this is consistent with my own observations of game compatibility in aggregate). So your (implicit) insistence that I claimed there are no blockers is obviously nonsensical already. But, furthermore, I went on to say this:
> unless it's been specifically engineered to block Linux
Which precisely describes a game that previously worked on Linux, with anticheat even, but now has been updated to remove that compatibility! It also describes a game whose anticheat has been configured that way since launch.
Also, your implication that "hard technical blocker" and "not worth it to implement" are interchangeable is ridiculous in any scope, whether you're Valve deciding whether to commit tens or hundreds of man-years to Proton development or a one man shop considering a hacky workaround in a client's dated web framework.
You are just barking at the mailman. I'm done here unless you can respond to the actual substance and nuance of my comments rather than a strawman.
Except we already know the anti-cheat system they use works on Linux, because other products use it. Therefore we have no evidence that the blocker here is technical: it could be political, process-driven, simple incompetence, or anything else. In any of the latter cases, there can be no conclusion of correctness of the parent comment's claim.
I don't play For Honour, but I'm monitoring Apex Legends pro scene and I know that Respawn has banned Linux because none of the existing anti-cheat solutions were good enough and the cost of maintaining their own solution was not worth the revenue Linux support was bringing.
Before the ban the cheater population was disproportionally high on Linux.
There's a segment of people who are into customizing their desktop environment as a hobby and end in itself.
Personally I've never really been into it, and these days I have a broad and revolving set of machines I have to use, so this sort of thing is absolutely not worth the bother. I just install KDE Plasma and use the computer.
I don't think this is true at all. If you don't have AI then either you build what you need to build and learn how to do it along the way, or you build nothing and learn nothing. There is no option to build something but learn nothing, but that's the default outcome with AI. To learn how to build something AI built for you, you need to build it again on your own, which I dare say very few AI users are doing for nontrivial projects.
reply