By the way, this is the same argument that Michael Burry used to short Nvidia.
He claims that GPU depreciation/obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He's being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0]
The logic is fundamentally flawed in my opinion. Let's use future Nvidia chips being much better optimized for LLMs for example.
New Nvidia chips 10x better than H200 --> data centers buy a lot --> Nvidia profits a lot.
New Nvidia chips 10x better than H200 --> data centers don't buy --> no faster than expected obsoletion.
In other words, the very act of buying many new Nvidia GPUs would be the event that causes faster than expected obsoletion. Yet, if you don't buy those new Nvidia GPUs, then there is no faster than expected obsoletion.
We also live in a world where there is competition. If Amazon doesn't buy but Microsoft does, suddenly Microsoft can offer better $/token prices.
1. The same isn’t necessarily true of the rest of the hardware stack which may be reused between accelerator generations.
2. You’re missing the “New Nvidia chips 10x B200, compute requirement grows less than 10*software improvements YoY -> buy less Nvidia.” Valuations are based on forward projections (>1T annual for NVDA) which can be revised down leading to a drop in valuation.
> If Amazon doesn't buy but Microsoft does
The big 3 all have their own proprietary accelerators. Meta is buying TPUs as well for now.
I would bet Nvidia’s major customers in 2 years are neoclouds and it seems that Jensen is making the same bet.
1. So this makes Burry’s argument even less convincing since those auxiliary hardware can last longer.
2. Jevons Paradox. More efficiency should lead to bigger models, faster inference, and more total tokens.
3. By all accounts, Trainium and Maia and Meta’s internal chip are struggling to keep up with Nvidia. That’s why they order as many Nvidia chips as possible. They’re not giving up but it isn’t as easy as buying stock Arm cores and taking them to TSMC.
Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.
1. Not really, current valuations are priced for persistent 80%+ margins based on spot. If auxiliary hardware lasts longer (I.e. next gen GPU reusing the same shell) then that reduces supply pressure and spot prices.
2. Jevon’s paradox is about total consumption, not margins. Valuations are about margins (and their projections). Many coal mine owners went bust despite increased total coal consumption.
3. Source? Gemini for example is 70% on TPU. I have yet to see data on Maia-300 beyond Microsoft PR. Remember it doesn’t have to be better it has to be more cost efficient. The overwhelming majority of inference spend does not care if token output is 20% slower if it is 50% cheaper.
> Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.
What Nvidia needs. Whether neoclouds can stay competitive vs hyperscalers paying Nvidia tax is far from clear, particularly when inference margins compress.
1. The whole Burry argument is that AI hardware becomes obsolete faster. If aux hardware can be reused, that works against the argument.
2. Total consumption drives more demand for the already supply constrained hardware. Can AI hardware market go bust? Sure it can. But being early is the same as being wrong in the investment market. When do you predict the bust to be?
3. Google, Amazon, Microsoft, Meta are all buying as many Nvidia GPUs as they possibly can. The biggest tell on how Nvidia is doing is that their share in inference has increased despite the increase in competition: https://archive.md/CKP0N. So while competition is getting bigger and bigger because the overall pie is getting exponentially bigger, Nvidia's growth is still higher than average.
> 1. The whole Burry argument is that AI hardware becomes obsolete faster. If aux hardware can be reused, that works against the argument.
Burry’s main argument is depreciation is being understated and the capex vintages will not be paid off before they are essentially useless. This can happen whether or not aux is reused.
> Total consumption drives more demand for the already supply constrained hardware.
Demand is the wrong metric.
Only number that matters is whether AI attributable revenue will be sufficient to pay back enough of each successive capex vintage (e.g. 750B this year, 1T next year, 1.2T in 2028) so that hyperscalers and neoclouds can either self-fund or continue to issue debt as bond markets are already straining and tax-payer backed sovereign debt is providing a high baseline. Otherwise they downgrade capex projections and the bubble pops.
Expensive compute needs expensive inference to justify 30-40B/year/GW of compute. There are many reasons why frontier API pricing which is what the industry is based on may not persist. It is also almost certainly the case that 2026 is the worst year of supply and demand mismatch to allow for 80%+ margins. HBF next year has the potential to single handedly pop the DRAM spot bubble.
> Can AI hardware market go bust? Sure it can.
This is the bear thesis. It is not that AI will crash or be useless.
> But being early is the same as being wrong in the investment market. When do you predict the bust to be?
Q4 27-Q2 28 is when the bill becomes due at the latest. There are sufficient financial levers left to buy time without returns until then.
> Google, Amazon, Microsoft, Meta are all buying as many Nvidia GPUs as they possibly can.
All of these companies have rock solid revenue streams and can easily swallow 500B of capex devaluation over time. Their buying of Nvidia today is not necessarily the indicator you are implying as there are strong competitive reasons to make the game more expensive for everyone else.
Burry’s main argument is depreciation is being understated and the capex vintages will not be paid off before they are essentially useless. This can happen whether or not aux is reused.
And why does he think depreciation is understated? It is because he thinks newer Nvidia GPUs will make older ones obsolete faster. Hence, my entire post.
The rest of your argument centers around whether AI growth will meet the cap ex expenses. I don't see anything new in it.
HBF next year has the potential to single handedly pop the DRAM spot bubble.
I'll believe it when I see it. Jevons paradox will apply here again in my opinion. HBF does not replace HBM.
He claims that GPU depreciation/obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He's being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0]
The logic is fundamentally flawed in my opinion. Let's use future Nvidia chips being much better optimized for LLMs for example.
New Nvidia chips 10x better than H200 --> data centers buy a lot --> Nvidia profits a lot.
New Nvidia chips 10x better than H200 --> data centers don't buy --> no faster than expected obsoletion.
In other words, the very act of buying many new Nvidia GPUs would be the event that causes faster than expected obsoletion. Yet, if you don't buy those new Nvidia GPUs, then there is no faster than expected obsoletion.
We also live in a world where there is competition. If Amazon doesn't buy but Microsoft does, suddenly Microsoft can offer better $/token prices.
[0]https://inferencex.semianalysis.com/inference