I had asked a question "how is Miami not on the list at all" and then realized that the table is actually an image not a table so I couldn't search it.
It's on the list. #34 out of 72, middle of the pack. Within the U.S. Chicago, NYC, Philadelphia, LA, and DC are above it. Houston, Dallas, San Antonio, and Phoenix are below it. They seem to weight heat-stress pretty highly in their rankings, hence why all the TX/AZ cities are at the bottom and are joined by South and Southeast Asia.
The list is a little weird in that I can't find any population-based reason why Miami would even be on it: it's almost the top-10 U.S. cities by population, but missing San Diego or any SF/BayArea city, while Miami & DC are #41 and #22 respectively. If you go by metro area it explains why Miami & DC are there and why Dallas & Fort Worth are condensed, but not why Atlanta is missing and San Antonio is included.
My guess is that Miami and Washington are there for essentially rhetorical reasons. Miami is often talked about as a city that's vulnerable to climate change, and Washington is the capital.
Have a single human AI chef. Everyone else has to write an engineering statement and submit it to the AI chef. All that interaction is outside the codebase. Engineers will take turns - perhaps 1 month stints - being the AI chef.
I guarantee you'll spend less on tokens, have better documentation, better code, and most importantly more competent engineers.
I actually quite like this idea and might try something like it on my team because we're defacto heading in this direction anyway and everyone's a bit frustrated, might be better if it was acknowledged and made official as something to try.
On the other hand, I think this denies the reality (in my experience anyway but I think enough people will agree) that one often solves a problem as they are working on it.
This method seems to presume that a good engineer will submit a well thought-out solution or direction giving the AI an extremely good overview of each problem and enough of a description of what to do that it will do things as expected and they can just review the result.
In my experience it just doesn't work that way in practice. One learns the problem and even the domain while developing the solution. So one would have to submit at least a half developed solution not just "instructions", for there to even be coherent instructions in the first place. And one needs that experience working on the problem to be able to properly evaluate a separately proposed solution.
All in all for me this leads more towards using AI as a co-developer than using it to just implement some fully thought out idea and then check what it did.
> This method seems to presume that a good engineer will submit a well thought-out solution or direction giving the AI an extremely good overview
Yes it does make that presumption - but that's part of the model here, that "prompt review" becomes the new code review and that the team is training on standard prompt semantics.
I'll share my outline - which is basically the one that was drilled into us in engineering school (CMU class 89)
1. Problem Statement
2. Definitions / Assumptions
3. Plan of attack / Options explored
4. Implemetation / Results
5. Limitations / Future Work
> So one would have to submit at least a half developed solution not just "instructions
The outline supports that in that you can fill out that part of the document which has been woeked on - including referencing any existing code.
Any other engineering students here recognize that outline?
There are most definitely is a moat - but it works both ways. The railguards in the models create moats keeping customers out. And the cost to build a modern agentic model is in the 10 figure range and growing. This is an expensive arms race that is going to create moats.
But most commodities are the same way. It’s super expensive to drill for oil. I need oil and I’m in no position to mine my own because of the massive capital investment. But it doesn’t stop it from being a pure commodity.
I couldn’t care less which company drilled for the oil… it’s all the same to me. Models are increasingly no different.
OpenAI and Anthropic are a gas station saying “buy our gas for 10x the price!” When the world is looking at them saying it’s just gas, we’ll take the cheaper brand. We’ve tested your gas and it’s really no better than the stuff that’s 1/10th the price.
Yes… and as the article says folks are still leaving some work to the big labs. But the big money is to be made at scale and those use cases don’t require OpenAI or Anthropic.
The crazy setup here is that even with that fraction of the pie these companies might be worth say $100 billion optimistically, which would be amazing in normal times. Problem is it’s a train wreck for their investors and the associated debt bubble if they can’t sustain a valuation of 1-2 trillion and the present setup does not put them on a course to that trajectory.
Of course. But that market won’t produce a $800 billion company, unless AI becomes ludicrously widespread — energy is used every day by virtually every person on the planet, and of course has plenty of mass consumption & “luxury” customers too.
I don't see any difference where I put gas from one place to another. But there is definitely differences between one model and another or even plans themselves .
Fair assessment. The challenge for OpenAI and Anthropic is that they need sizeable margins to pay for the massive costs incurred. Market forces are driving things in the opposite direction and fast.
When your competition has a tiny cost base compared to yours and lacks the bonkers future capital commits you made then that’s a terrible position to be in… hence their conundrum.
With other digital goods the distribution and operating costs have been essentially free. No business worried that much about the cost of running Microsoft Office on the PCs they already distributed to their employees. They were only concerned about the licensing costs. And Microsoft didn't worry about the cost of printing CDs or the costs of serving Office online. It wasn't zero, but again negligible compared to the cost of development and the licensing costs.
For LLMs the costs of training and inference are a very significant part of the overall costs.
Except OpenAI and Anthropic has brought in a lot of money, that with this trajectory will make it some of the worst investments in ”software” ever (if it’s true enterprise clients are actively moving away, I know we are but for other reasons).
But they're all converging on capability. Do I care if it's a 72% or 74% on SWEBench? Practically, probably not. And if I'm not paying per token locally, then if it takes a tiny bit longer to get to the result, I don't care.
Smartphones and laptops are also "converging", but Apple is always a year or two ahead so it doesn't matter. "Converging" is a meaningless term when things move quickly and cost billions to develop.
1. People buy Apple because of the broader ecosystem of products and the “it just works” aspect of that ecosystem. Other companies make phones with features that are objectively better but folks don’t switch because the Apple ecosystem is sticky.
Despite trying, neither OpenAI nor Anthropic has managed to move up the stack beyond “hey guys new model release today!” announcements that everyone yawns at.
2. Switching costs are real. It’s a PITA to switch not just the phone but everything else. Switching model providers is a line of code and takes almost no effort.
Apple has a true moat which is why they can command a premium. OpenAI and Anthropic have no moat which is why they’re in trouble.
Apple would have probably lost more customers if it didn't also create a hard-to-leave lock-in ecosystem.
Lots of people use old iPhones and don't care about some "up to" benchmark bumped every year, but are stuck with iMessage contacts, their stuff in iCloud, Apple Watch or apps that are not allowed by Apple to even mention they have Android versions.
Honestly not. Most big corporates have arrangements were all the major models and now open models are available from the same API endpoint. It is literally one line of code to edit in most cases, even more so in big companies.
I don't think most motorists would care if OpenAI's gas stations just released 106-Octane "Intersteller" gas, unless their cars specifically require it.
Even commodity markets recognize different grades of product. The oil market separately prices different grades of oil, different refined products. All that really matters is that when you go to the market to buy, you can say, "I need X amount of this grade of this product" and that is what you will get. If AI models can be sold that way, you basically have a commodity market.
That's true for sure in most business endeavors. The goal is to fill the area under the demand curve and there are demands for F1 race cars and for scooters. The analogy breaks down somewhat with software in general and for sure with superintelligence. A superintelligence can provide those "low-level" (ie scooter) services perhaps just as effectively because it's super intelligent and knows how to do things efficiently - for example by spawning agents of different intelligence levels. It can thus fill the area under the demand curve. This is what the big AI firms are shooting for.
seems self evident if you read the new. You can Google it yourself, but here's the results from my googling - and this is just for the hardware. Double that to add personnel and corporate infrastructure
"To build or purchase the physical hardware required to store tens of petabytes of data and train a State-of-the-Art (SOTA) frontier AI model, you are looking at a capital expenditure (CapEx) ranging from $320 million to well over $1 billion."
No, it’s not. Rumsfeld segregated what we know from what we know we know (and vice versa). An unknown (whether known or unknown) may be knowable or unknowable—his framework doesn’t address knowability.
> because an LLM is not a God. It is not an unknowable
I tend to agree with you. This has nothing to do with the Rumsfeld comparison being wrong.
reply