Hacker Newsnew | past | comments | ask | show | jobs | submit | epolanski's commentslogin

If model X fits your need, you don't need to upgrade.

I have released applications on Gemini 3.5 flash that make real money and I don't see any particular reason to upgrade.


It feels cheap plastic, true, but it's also light and resists a lot of damage.

Mine v10 has 6 years, has seen some important falls, but bears no damage of any of those.

I wouldn't recommend it because price/quality Dyson vacuums feel like you're insanely over paying.

I've bought my grandma a 150$ Levoit LVAC-200-WEU, and I have used it a lot. I swear it lasts more than the v10 and it cleans better. It also feels much higher quality to touch.

It's cons are that it didn't come with a narrow tip to suck odd corners and that it's slightly heavier than the v10. But it costs also less than half and I see no victory at all for the Dyson product.


I think this whole distillation argument is between fully overblown and bogus.

In any case, highly misunderstood.


Interesting, I liked to experiment with a second model "simplifying" and summarizing the previous messages and continue.

Needless to say, it improved output on following messages by whatever metric I cared for.

Not sure why would they prevent it.

I give you a chain of messages, what do you care for what the origin is?


While I also agree that Opus 4.6, in some ways, was the last model that truly felt an assistant, all the following ones seem to have inverted the role, even a blind person can see that throwing difficult problems, and complex bugs at this model achieves more than predecessors.

I don't think there's nothing ground breaking, but sure it achieves and finds more, sooner.


Just the other way I was thinking that if I asked "What does Lamborghini do?" the only correct way to answer is a single sentence "Which Lamborghini are you referring to?".

But LLMs will fail at this question: they will tell you about Lamborghini's latest car and mix some history in it. Just try.

Which is the wrong answer anyway, because there's at least two major companies called Lamborghini, one making cars, one making agricultural equipment and at least one famous person (Elettra) with that family name.

This very simple test/question makes me realize how much do I hate LLMs in a sense: while I agree that the answer it gives is the most plausible for 90% of the users, it's ultimately both wrong and long. And that 90% compounds.

But there's no "correct" answer in my eyes than "who are you referring to?". Possibly without listing all the possible Lamborghinis.


This is ... unnecessarily pedantic. Anybody in my social universe who asked me that question would undoubtedly expect "they make cars".

If you're picking nits, why not focus on the word "do" and (wrongly) expect an answer like "Lamborghini (either of the two main companies of that name) does not 'do' anything - the companies employ humans who 'do' things. Lamborghini is a legal entity established to allow humans to 'do' things, such as make cars, or agricultural equipment."

Shared context is a thing. Reducing every conversation to first principles is not always required. Get a grip.


This is not a nit. This is a real problem in the technology.

It assumes the average and plausible answer token by token.

And this tendency shows in every single field it's applied to.

At the end of the day I want *correct* answers, to the point.

Instead LLMs, no matter if it's version 3500, are bound to producing average results: slop.


> As someone working in science, this belief confuses me. How (by what means) do you think Fable 5.1 will be able to make further progress in scientific domains?

The same way it did in the previous versions: brute force.

I don't believe that LLMs have any particular intelligence we don't, but there's an endless list of problems we either don't have bodies to throw at, or the bodies we can throw at it, don't have such a huge large context to crunch problems.

What LLMs will always intrinsically fail at is showing us genuine new intuitions. The technology is about predicting the next plausible token/sentence.

They will not revolutionize human knowledge, but they can definitely widen it a lot.


> They will not revolutionize human knowledge, but they can definitely widen it a lot.

I am generally quite enthusiastic about all this, but my biggest fear is that we will not recognize the extreme need for more scientists at a time when there is so much more science to be done. The rate of scientific understanding must keep pace with the amount of science being output, both for verification and further discovery. It's a pipelining issue, and I predict a stall in the bits that require the (currently rare) people who know what they're doing.


We are not limited by intelligence, or bodies.

Why would you think we aren't?

There's an endless number of scientific problems out there in any field, and nobody able to dedicate themselves to it.

I've been in research (you con check my name on Google Scholar for my released papers), there was always an endless number of experiments or paths more I could've taken than the time and resources to do so.


In many fields the limitation is not thinking. In my field (particularly obscure UHV surface science) we are limited by experimental results, and that experimental data is limited by the number of operable machines in particular configurations. These are multi-million dollar specialized machines that are artisanally made. There's a small, single-digit number produced each year, and each one is hand-calibrated to its task.

Due to computational limitations, this is not work that can be effectively simulated on a classical computer. Actual experimentation is required.

I fail to see what impact improved AI would have on this problem. Perhaps better selection of experimental problems for our limited capacity to run experiments, but that's assuming there is any slack left to take up. In reality we already have more brainpower than needed applied to this problem.


> OpenAI and Anthropic have totally dropped the ball on getting their respective desktop apps in front of the enterprise business user cleanly

Data, contracts, procurement are the blockers for enterprise, not features.

Big companies will use whatever AI tools Google or Microsoft because they are already on the Microsoft/Google suite.

No amount of shiny new tool can compensate here, by the time somebody to buys it, months pass and those two companies will have it anyway.


You highly underestimate how both Claude and OpenAI have peanuts of the enterprise marketshare compared to Copilot and Gemini.

I see it first had across all my non-tech friends: their companies already used Teams/Sharepoint or Google Suite. Those added AI capabilities with some minimal vetting/setting by the org. Data retention and contracts, the hard parts, were already handled because those are new features/extensions of the same products they had.

Comments like yours seem to be screaming "HN bubble". The real world doesn't care and will wait for Microsoft/Google to offer the same stuff, hell, even HN apparently barely knew what Claude and OpenAI work offerings did till today.


+1, I see it very clearly at my SO work.

Half of her organization has just a calendar filled with meetings.

And without meeting the organization would find that you only really need a third of the people, and you would even likely increase the overall output.

Many time wastes are just designed to make people busy, not productive.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: