Zero Times Infinity

Audience: Engineers, operators, and investors trying to tell which AI infrastructure costs are falling, which are compounding, and who is on the hook when the bill comes due.
Reading time: ~17 minutes.

Two things I keep hearing about AI, often from the same person in the same conversation. The first is that it’s absurdly expensive: the datacenters, the gigawatts, the Nvidia market cap, the electric bill. The second is that it’s basically free now. A GPT-4-class answer that cost $30 per million tokens in 2023 costs a nickel today, and a frontier-class open model runs on a secondhand Mac in my office. Most of the confusion about AI economics comes from treating those as a contradiction. They aren’t: one is a price, the other is a bill. A bill is a price times a quantity, and the quantity is growing faster than the price is falling. The price is heading toward zero and the quantity toward infinity, and zero times infinity is not a number you can look up. Economists have a name for the pattern, Jevons’ paradox, from the observation in 1865 that more efficient steam engines made Britain burn more coal rather than less. What’s new is the number of zeros.

Four companies will spend about three quarters of a trillion dollars on AI infrastructure this year. Nvidia is 8% of the S&P 500. The memory in your laptop has roughly tripled in price since last fall. Google’s models process more than three quadrillion tokens a month. And in June the Bank for International Settlements put “AI capex bust” on its short list of things that could break the global financial system.

So this essay is about the bill: where it goes, who’s paying it, and how the two curves resolve. I’ll follow a dollar from a token back to a wafer, and on the way check the three intuitions I hear most often: that AI is driving up everyone’s electric bill, that the Nvidia-AMD-OpenAI-Oracle money-go-round has to end in a crash, and that there’s something called a RAM supercycle. All three are half right.

What a Gigawatt Buys

A year ago people counted AI in GPUs. Now the cool people in tech count it in gigawatts.

A gigawatt is a billion watts, continuously: the average draw of about 800,000 American homes, and now the standard size of a serious AI campus. What one costs depends on who you ask. Jensen Huang, who has a certain interest in the number, said $50 to $60 billion in August 2025 and $80 to $100 billion this June for the next generation; Nvidia’s own accounting is that it collects about $25 billion per gigawatt of Blackwell and $40 billion per gigawatt of Vera Rubin, which is most of why the headline went up. Epoch AI, a research group that tracks the industry’s numbers, built the most transparent model I’ve seen: a one-gigawatt Blackwell facility at $38 billion, of which $21 billion is servers, $11 billion is the building with its power and cooling, $5 billion is networking, and about $300 million is land and substations. Run it for a year and the fully loaded cost is around $8.5 billion, 60% of it the servers being written off.

Where the money goes in a one-gigawatt AI datacenter (Epoch AI model) One gigawatt, Blackwell generation (Epoch AI, May 2026) Build: $38B servers 56% building 30% net 13% Per year: $8.5B server depreciation 60% facility 16% net 14% electricity 7% Servers on a 5-year life, facility on 14. Land and substations are the 1% sliver at the far right of the top bar.

The building is real estate. The networking is plumbing. The chips are the money, and the chips are bought from one company at a 75% gross margin; by one Wall Street estimate, Nvidia’s gross profit alone is about 29% of everything anyone spends on an AI datacenter. Behind Nvidia sit TSMC, which fabs and packages every one of those chips, and the three memory makers who supply the high-bandwidth memory stacked on top. A gigawatt is a purchase order that flows to Santa Clara, then Hsinchu, then Icheon and Boise.

OpenAI’s chip deals alone come to about 26 gigawatts, plus 4.5 from Oracle, plus Sam Altman’s own framing of “about $1.4 trillion over the next eight years.” Anthropic’s commitments add up to something north of 13 gigawatts, with overlaps; Meta’s Hyperion site in Louisiana is designed for 5. At $40 to $60 billion a gigawatt, these are the trillions you keep reading about, and they are arithmetic on letters of intent. What’s energized is a different number. As of April, of Stargate’s seven US sites, only Abilene, Texas had power flowing, at most 0.6 gigawatts; the other six are foundations and steel due in late 2028, and the Abilene expansion was abandoned in March. Sightline Climate found nearly half the US capacity slated for 2026 delayed or cancelled, mostly over power, transformers, and neighbors. The first single-tenant sites reached roughly a gigawatt this summer, and Epoch measures the largest one doubling every ten months.

So the announced number is tens of gigawatts and the energized number is a few, and the gap is a power story. Chips are tight too; TSMC’s CEO said in July that packaging capacity “limits my customers’ growth.” But a chip queue is measured in quarters. A heavy-duty gas turbine ordered from GE Vernova today arrives in 2031.

The Power Bill

US datacenters used 4.7% of the country’s electricity in 2024, and Lawrence Berkeley National Laboratory’s central case for 2030 is 11.8%, a third of all US load growth this decade. Those numbers are real, and they’re where the intuition that AI is a power hog comes from. Per query, the intuition is off by orders of magnitude: Google measured the median Gemini prompt at 0.24 watt-hours, about nine seconds of television, and even a heavy reasoning query with 100,000 tokens of context runs to 40 watt-hours, less than two minutes of a hair dryer. The aggregate is right because the query count is in the quadrillions a month.

Here’s the part that changed how I think about all of this: electricity is not where the money goes. In Epoch’s model, the annual power bill is $594 million against $8.5 billion of total cost, 7%. If power were free, tokens would get about 7% cheaper. If the GPUs were free, they’d get 60% cheaper. That asymmetry explains behavior that otherwise looks insane: xAI running dozens of unpermitted gas turbines outside Memphis, Microsoft prepaying for substations. Nobody in this business is optimizing the power bill. They’re optimizing the number of days $20 billion of silicon sits in a crate. Satya Nadella said it plainly last November: “It’s not a supply issue of chips; it’s actually the fact that I don’t have warm shells to plug into.”

That’s also why “AI is raising my electric bill” is a fight about allocation rather than consumption. In PJM, the grid from Chicago to Washington, the capacity auction went from $29 per megawatt-day in 2024 to a regulatory cap around $330, where it has been pinned for three straight auctions, and the market monitor attributes $29 billion of the $64 billion those auctions collect to datacenter load. Nationally, though, the last five years of rate increases trace mostly to grid investment, wildfires, and gas prices, and Texas and Virginia, the two biggest datacenter states, had among the smallest. What decides the next five years is whether the people building gigawatts are made to pay for the turbines and transmission built to serve them. They can afford it; power is 7% of their cost. Whether they’re made to is politics, and the politics arrived this year: more than 300 datacenter bills in 30-plus states, large-load tariffs in 23 of them, fifteen gubernatorial candidates running on moratoria, and a Texas freeze on new grid connections in August after its interconnection queue hit 474 gigawatts. Wood Mackenzie reckons about 72% of US datacenter power requests are phantom.

And power is what binds the whole buildout. GE Vernova’s turbine backlog went from 100 to 116 gigawatts in a single quarter; the nuclear deals of 2024 and 2025 are mostly still paper, and if every one were built it would cover less than a fifth of projected 2035 datacenter demand. China added around 430 gigawatts of wind and solar in 2025 alone. The US added 53 gigawatts of everything.

The Memory Tax

The transfer from consumers to the buildout that actually happened in 2026 didn’t come through the meter. It came through memory.

The mechanism is a wafer trade. A gigabyte of high-bandwidth memory, the kind stacked next to every AI accelerator, takes three to four times the wafer area of ordinary DRAM, and AI will consume about a fifth of the world’s DRAM wafer output this year. The fabs chose the datacenter, and did it without adding wafer starts, because a new memory fab takes until 2028. Conventional DRAM contract prices rose a record 90% to 95% in the first quarter of 2026, then another 60 in the second; J.P. Morgan puts the cumulative rise since the start of 2024 above 400%. The DRAM industry’s quarterly revenue went from $27 billion in early 2025 to $155 billion this spring. SK Hynix reported a 76% operating margin for the June quarter, a point above Nvidia’s gross margin, and Samsung’s chip division expects to earn more this year than in its previous forty combined.

Everyone downstream paid. In June Apple raised Macs and iPads by $100 to $300, saying “we have never seen a component price increase this much, this quickly,” and in September it raised the price of every iPhone it sells by $100, last year’s models included, same hardware. The console makers, Amazon’s devices, and Nvidia’s own desktop AI box went the same way; even Nvidia can’t get enough memory for Nvidia. Gartner has PC prices up 17% this year and shipments down 10. The Minneapolis Fed attributes about 0.4 points of core inflation to AI hardware demand, an effect it calls at least as large as tariffs. And the squeeze runs up the chain: memory was about 9% of the value of a Blackwell rack and is about a quarter of a Vera Rubin rack, Nvidia guided its gross margin down to 71% or 72%, and it told customers that systems shipping in early 2027 will cost 15% more. Three memory makers are now taking margin from the company that takes margin from everyone else.

In July the memory stocks crashed, 30% to 50% from their peaks, on fears that HBM expansion and a new Chinese entrant would swamp the market in 2027, in the same month Microsoft and Amazon raised capex guidance and blamed component inflation. Both were right; contract prices lag spot, and equities discount next year. Contract prices are still rising, at low-teens percent a quarter rather than 90, and the brokers who model it have supply and demand balancing in 2028.

I’ve written about the local-hardware side in Agentic AI @ 2 FPS: in this cycle, unlike 1998, the datacenter outbids the consumer for the wafer. If you want to know what AI cost ordinary people in 2026, don’t look at the electric bill. Look at the price of a laptop.

The Circle

On the last night of 1999 I was on Microsoft’s Redmond campus with a walkie-talkie. I ran a development team on MSN’s accounts and billing systems, and the radio net existed for the scenario where everything else failed at midnight. Nothing did. The other time zones had already rolled over, so by the time Pacific midnight arrived we were basically partying, and since you couldn’t tell who was talking, people were cracking jokes on a channel that had vice presidents on it. Possibly the vice presidents were partying too. There was no runbook to speak of; those were the wild-west years of online distributed systems, and next to what I’d see at Amazon sixteen years later it was comically naive, though par for the course. The stock market never came up. It was engineers running the show, and the financing of the thing we were standing in was somebody else’s department. Ten weeks later the Nasdaq peaked, two weeks after that Cisco passed Microsoft as the most valuable company on earth, and by the end of March I’d left for a startup. It took me years to see the lesson: the people running the infrastructure never see the financial structure they’re inside. The circle gets drawn in a different room.

Here’s what that room looks like today: a circle, in which the chipmaker funds the customer, the customer’s cloud providers borrow to build for the customer, and everyone buys from the chipmaker. The deals first, because the shape is the argument.

In September 2025, Nvidia announced it would invest “up to $100 billion” in OpenAI as OpenAI deployed 10 gigawatts of Nvidia systems. It was a letter of intent; five months later no contract had been signed and no money had moved, and in March Huang said the $100 billion was “probably not in the cards.” What replaced it: $30 billion of straight equity in OpenAI’s $122 billion round at an $852 billion valuation, and then, in August, a guarantee capped at $105 billion backstopping OpenAI’s twenty-year leases on 4.25 gigawatts of datacenter at a SoftBank-owned campus in Ohio that, per Nvidia’s filing, “will exclusively host NVIDIA AI infrastructure.” Nvidia backstops the rent on a building that will only ever hold Nvidia chips, bought by a tenant Nvidia owns a piece of.

AMD signed OpenAI to 6 gigawatts of GPUs and issued it warrants for about 10% of the company, vesting as the gigawatts deploy, then signed the identical structure with Meta; two customers can now earn a fifth of AMD by buying AMD’s product. OpenAI’s largest single commitment is about $300 billion of cloud capacity from Oracle over five years, and Oracle is borrowing to build it: $43 billion of debt last fiscal year, another $40 billion of financing planned for this one, capex of $90 to $95 billion against operating cash flow of $32 billion, free cash flow of negative $23.7 billion. Most of that goes to Nvidia. OpenAI is about half of Oracle’s backlog; S&P cut Oracle to one notch above junk in July. Microsoft and Amazon close the loop from the other side, as the diagram shows, and Anthropic’s version has the same shape with Nvidia, Microsoft, Amazon, and Google. And in January Cerebras, a chip startup, took a billion-dollar working-capital loan to build capacity for OpenAI. The lender was OpenAI. A company that expects to burn $25 billion this year is financing its own supplier, which is 1999’s vendor financing run in reverse.

The circle: who funds OpenAI, what OpenAI owes them, and where the money goes next ...and buys Nvidia ...and buys Nvidia OpenAI valued at $852B, burns ~$25B/yr Nvidia $30B equity $105B lease guarantee orders 10 GW of systems AMD warrants: ~10% of AMD orders 6 GW Amazon $50B in, $138B AWS out Microsoft owns 27%, $250B Azure Oracle $300B contract, rated BBB- $300B over 5 yrs (Oracle borrows to build) green: money into OpenAI    blue: what OpenAI owes back    orange: where it goes next Anthropic's loop with Microsoft, Nvidia, Amazon and Google has the same shape.

Nvidia’s balance sheet is the clearest record of how far this has gone. Its equity investments went from $2.2 billion in 2024 to $99 billion in July, about $50 billion of it in the frontier labs. Guarantees total $108.5 billion. There is a $36 billion commitment to buy cloud capacity back from the AI clouds Nvidia sells chips to, which the filing calls a “new business model” and explains exists because those clouds “lack the ability to secure long-term infrastructure contracts and investment-grade financing capacity.” And in August Nvidia signed up Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR for financing platforms meant to mobilize $500 billion of other people’s capital, with Nvidia backing up to 25% of each deal. Huang: “In AI, compute is revenue.” Asked about the word circular, he called it “ridiculous.”

Is it circular? Partly. Nvidia’s cash into companies that buy Nvidia chips, set against about $300 billion of trailing revenue, is a single-digit to mid-teens percentage even if every dollar came straight back; Michael Burry’s “100% circular” is about forward commitments, not recognized revenue. The money flows one way, the way GM Financial’s does, so the deals don’t inflate today’s sales. What they do is correlate the risk. Every balance sheet in the loop now keys off the same variable, whether OpenAI and Anthropic keep raising money, and the contingent layer of guarantees and financing platforms is several times larger than the equity.

The precedent everyone reaches for is Lucent, which in 1999 had $8.1 billion of vendor financing outstanding on $38 billion of revenue. Forty-seven of those customers went bankrupt, Lucent’s revenue fell 69% in three years, and telecom bondholders recovered about twenty cents on the dollar. The difference in 2026, the one the bulls are right about, is that the customers have cash flow: Microsoft, Alphabet, Amazon, and Meta generated $451 billion of it in 2024, and the carriers of 1999 generated none. That difference is shrinking. The big four now spend roughly all of their operating cash flow on capex. Meta’s free cash flow fell 91% in the second quarter and it stopped buying back stock; Alphabet’s went negative and it raised $85 billion of equity in June, the first hyperscaler to sell shares for AI. Hyperscalers issued $182 billion of investment-grade bonds by mid-July, about 15% of all US corporate issuance. The “production capital, not financial capital” argument that made this feel different from 1999 was true in January. It’s fraying by the quarter.

Who Holds the Bag

So if the circle breaks, who’s holding what? Sort by tier, because the exposure is uneven and the credit market has already done the sorting.

Tier Who Exposure if the music stops
Hyperscaler equity Microsoft, Alphabet, Amazon, Meta Mild. Debt-to-equity in the single digits to low twenties, credit default swaps around 49 basis points. A bust is a capex cut and an earnings dip as depreciation catches up.
Intermediaries Oracle, CoreWeave and the neoclouds, private credit and SPVs Severe. Oracle at BBB- with negative free cash flow and half its backlog from one customer. CoreWeave with $35 billion of debt, paying $640 million of interest a quarter against $128 million of adjusted operating income, swaps near 855 basis points, roughly a coin flip on default within five years.
The labs OpenAI, Anthropic, xAI The bet itself. Equity-funded, about $190 billion raised between them in eighteen months at valuations near a trillion. OpenAI projects burning $25 billion this year and $57 billion next against a compute plan north of $650 billion through 2030.
Sovereigns and index funds Gulf funds, SoftBank, your 401(k) Structural. Nvidia is 8% of the S&P 500; the seven largest stocks peaked at 35% of it in June. Passive holders own this concentration by construction.

The middle tier is where 1999 lives, because that’s where the debt is. Meta’s Hyperion campus, for instance, belongs to a joint venture with Blue Owl, a private credit firm, that sold $27 billion of bonds; the bonds are rated A+ because Meta guarantees their residual value, and that guarantee doesn’t show up as debt on Meta’s books. Multiply that structure across the buildout: Morgan Stanley estimates $1.5 trillion of it through 2028 has to come from outside the hyperscalers’ cash flows, more than half from private credit, and the BIS warned in June about terms “poorly disclosed, with risks of the same asset being pledged multiple times.” The stress is already visible at the weakest link. In July CoreWeave’s lenders repriced a $2.6 billion loan and added covenants, and its stock fell 36% that month; Jim Chanos calls the neoclouds “financial conduits, not technology companies.”

The labs are the layer everyone else is lending against, which makes their ability to keep raising money the load-bearing variable in the whole structure, and the market knows it. When OpenAI’s CFO said in November that she’d like a federal “backstop” for AI infrastructure financing, she and Altman walked it back within 48 hours, and Nvidia still fell 7% that week. The closest thing to a real backstop so far is private: xAI’s 12.5% junk debt became investment grade in June by being folded into SpaceX and refinanced against Starlink’s cash flows.

Here’s the sort the credit market has made, in one table.

Borrower Five-year credit default swap, mid-2026
Amazon, Alphabet, Microsoft (composite) ~49 basis points, highest since 2018
Oracle ~200 basis points, an 18-year high, rated BBB-
CoreWeave ~855 basis points

The market thinks the risk sits with the intermediaries, not the hyperscalers or Nvidia. I think the market is right, with one caveat: the intermediaries are load-bearing. Oracle’s $300 billion, CoreWeave’s $22 billion, and the Ohio campus Nvidia is guaranteeing are a large share of the capacity OpenAI is counting on, and OpenAI is the customer Nvidia’s own filing describes as “a meaningful amount of our revenue” through those intermediaries. A default in tier two doesn’t stay in tier two.

The Price of the Frontier

The other half of the bill is what gets run on the capacity: training the models and serving them. Training first, because it’s the part with the most mythology.

The cost of the biggest training runs has grown about 2.4 times a year for a decade, by Epoch’s count. GPT-4’s final run in 2023 cost roughly $40 million in amortized hardware and power; the most expensive publicly estimated run to date is Grok 4, in mid-2025, at about $490 million and 310 gigawatt-hours. Dario Amodei’s ladder was $100 million, then $1 billion, then $10 billion by 2026, and as far as anyone can document no single run has crossed a billion; the $10 billion figures are program costs. Those are the real number anyway: the final run is about a tenth of a frontier lab’s research compute, and OpenAI’s research compute is projected at $19 billion this year. That’s also the right lens on DeepSeek’s famous $5.6 million: true, and the final pretraining run only, on a fleet SemiAnalysis costed at about $1.6 billion. Cheap final runs on expensive fleets are how the whole industry works; DeepSeek just said the number out loud.

Two shifts in 2026 change what “training” means. Post-training overtook pretraining: GPT-5 used less pretraining compute than GPT-4.5 while scoring higher, because the work moved to reinforcement learning. And the fast-follower discount got measured. Epoch puts open-weight models an average of four months behind the closed frontier, on an index where the frontier moves fourteen points a year, and the closed frontier trains on ten to twenty times the compute of the open one. Distillation is how: Anthropic reported one Chinese lab extracted more than thirteen million exchanges from Claude to train its own model.

Put those together and a frontier model is a depreciating asset with roughly a one-year half-life. A lab spends a billion-dollar program to hold a lead the commons matches in months at a tenth the cost. That’s why the labs race, why the training bill keeps rising even as last year’s intelligence goes free, and why the money has to be made during the months a model is ahead. Which is the serving business.

The Token Machine

The commodity tier did exactly what everyone predicted. GPT-4 launched in March 2023 at $30 per million input tokens; GPT-4o mini, which matches or beats it on most standard benchmarks, was $0.15 in mid-2024; GPT-5 nano is $0.05. That’s a 600-fold decline in thirty months at a fixed level of capability, the “tenfold a year” that a16z called LLMflation.

Price down, frontier price back up, volume up: 2023 to 2026, log scales Tokens: commodity price, frontier price, and volume (log scales) 2023 2024 2025 2026 $30 $3 $0.30 $0.03 per M input 10Q 3.2Q 1Q 100T 10T per month GPT-4, $30 GPT-5 nano, $0.05 cheapest GPT-4-class Anthropic flagship list price Opus 4.5, $5 Fable 5, $10 9.7T tokens/mo Google tokens per month Commodity price: cheapest OpenAI model that matches the original GPT-4 on standard benchmarks, input price. Frontier price: Anthropic's top model, input. Volume: Alphabet's disclosed monthly token counts.

But 2026 added a second curve, and it’s the one I’d want a board to understand: the price of the frontier went up. OpenAI’s flagship was $1.25 per million input tokens when GPT-5 launched in August 2025; the current GPT-6 Astra is $10. Anthropic went from $5 for Opus 4.5 to $10 for Fable 5. Google has announced its Flash prices double on January 1, 2027. Anthropic’s newer tokenizer produces about 30% more tokens for the same text, a price increase that never appears on a price page. So there are two curves now. Last year’s intelligence still gets ten times cheaper every year. This year’s got more expensive, because the labs discovered they have pricing power at the frontier, during the exact window the previous section says they have to monetize.

Volume is the other half of the identity, and it surprised everyone, including the people selling it. Google processed 9.7 trillion tokens a month in May 2024 and 3.2 quadrillion in May 2026, a 330-fold increase in two years. Epoch estimates token demand growing about ten times a year against inference capacity growing three to four times, which is a crunch by definition. Reasoning models generate around four times the tokens per query and now carry more than half of all routed tokens, and an agentic coding session is millions of tokens that run unattended.

Coding is where the volume comes from. Claude Code went from a $1 billion run-rate last November to roughly $8 billion by May; Anthropic’s total went from $9 billion at New Year to $65 billion at the end of July, and Amodei’s line was “we tried to plan very well for a world of 10x growth per year. And yet we saw 80x.” Coding leads because its return is legible: a verification loop turns tokens into merged, tested code, and a team can see the conversion rate. I’ve argued in The Bitter Lesson of Agentic Coding that the quality of that loop is the ceiling on what autonomous coding loops can produce; it’s also the reason the tokens get bought. Ramp’s card data shows the median company spending about $12 per employee per month on AI and the top 1% spending about $7,400. That 600-fold spread is the demand story in one number. Most companies are dabbling. A few have found a loop that converts tokens into money and are feeding it as fast as their vendors will let them.

The “sold below cost” worry, which I shared, turns out to be mostly wrong at the labs and right at the resellers. The Information reported OpenAI’s margin on paid inference at about 70% as of last October, double its early-2024 level, and SemiAnalysis puts Anthropic’s inference margin in the mid-60s, up from 38% in 2025 and negative 94% in 2024; blended gross margins are lower because the free tier and the research sit in the same line. Paid tokens aren’t underpriced. Where inference really was sold below cost was one layer up: Cursor ran a negative 23% gross margin in the quarter ending in January, per The Information, repriced twice, and got acquired by SpaceX in June. And an H100 that rented for under $2 an hour at the end of 2025 rents for about $3.40 today.

So here’s the bull case in a single line: revenue is price times volume. For revenue to grow, volume has to outrun the price decline, and it has, by a wide margin. Every dollar of that revenue is somebody’s evidence that the capacity will be used.

How This Ends

Start with the historical shape: this is an installation period, in Carlota Perez’s sense: financial capital overbuilds the infrastructure, a crash transfers the assets to production capital at a discount, and the productive deployment happens afterward on infrastructure someone else paid for. Every analogy people reach for fits the pattern and was proportionally larger. Britain’s railway mania put 7% to 8% of GDP into rails in 1847 and a third of the authorized lines were never built. The telecom buildout put more than $500 billion into fiber, about 1.2% of GDP at the peak; 97% of it was dark in 2002, bondholders got twenty cents, and then the dark fiber became YouTube and Netflix and the cloud. AI capex is running at roughly 1.2% to 1.5% of US GDP. Jeff Bezos called this an “industrial bubble,” as opposed to a financial one, and meant it as a compliment: the industrial kind leaves things behind.

The twist is asset life, and it cuts both ways. Fiber lasts thirty years. A GPU is booked over five or six and is frontier-competitive for one to three; Burry’s claim that the hyperscalers are understating depreciation by $176 billion is a fight about whether the right number is three years or six, and the evidence is mixed. Either way a GPU glut clears in a couple of years; it isn’t the decade of free capacity dark fiber turned out to be. But look at what else is in the $38 billion: the building, which Microsoft now depreciates over 25 years, the substations and transmission, the turbines ordered for 2031, the fabs and packaging lines. Those outlast any correction. The durable legacy of an AI bust would be power and fabs, plus a fire sale of compute.

The consolation, if you build on this stuff rather than finance it, is that intelligence gets cheaper in every branch. In the boom branch, tokens get cheaper on schedule: chip performance per dollar compounds at 49% a year and each hardware generation brings a two- to ten-fold drop in cost per token. In the bust branch, they get cheaper faster: stranded GPUs, stranded power, labs selling capacity at marginal cost to service debt. The bubble question decides who ends up owning the capacity. The capacity exists either way. Plan for abundance.

What I’m watching is the ratio, because it says when. Capex is a bet that volume outruns price for long enough for revenue to catch it, and on the back of my envelope revenue catches capex around 2029:

Mid-2026 Growth rate
Identifiable AI revenue, run-rate (OpenAI, Anthropic, Microsoft AI, Amazon AI, with overlaps) ~$150-200B 2-3x per year
Big-four hyperscaler capex, 2026 ~$750B ~1.7x per year
Required end-customer revenue by Sequoia’s rule of thumb (4x Nvidia’s run-rate) ~$1.2T

That is roughly where the careful forecasters land too; Goldman doesn’t see supply and demand balancing before 2028. So the dangerous window is the next two years, while capex is still compounding and revenue hasn’t caught it. Five things will tell you whether the window is closing badly: OpenAI’s next funding round, Oracle’s credit spreads, whether Stargate’s remaining sites energize on their late-2028 dates, one-year H100 contract prices (the best anti-bust datapoint right now), and token volume, which has to keep growing faster than the commodity price falls. The day it doesn’t, revenue stops growing and the whole tower is built on a flat line.

I don’t think that day is close. I do think the middle tier of the capital stack is going to have a very bad year before 2030, and that when it does, the GPUs will change hands at a discount, the power will still be there, and the people who bought the discount will run the deployment period on it.

Right but Early

I’ve been on the wrong side of that timing once. The startup I left Microsoft for in March 2000 was Wildseed, built on the idea that phone software was being written by radio engineers, that the radio stack and the operating system would separate into layers the way PC hardware and Windows had, and that the app explosion of the PC era was about to happen on phones. We were describing the mobile revolution seven years before the iPhone and eight before Android, and the wireless industry was blind to it. We were also a software company trying to ship hardware into the capital drought that followed the crash, which is a bad combination. We built the phone in Korea and sold it on second-tier carriers; it was genuinely cool, it was not a market success, and AOL bought the company in 2005. The thesis was right. The deployment period arrived on its own schedule, on networks the telecom bubble had overbuilt and the crash had repriced, and it arrived for other people. That’s the shape I expect here. The last GPU buildout I was part of, in 1998, was financed by teenagers buying Quake. This one is financed by Blue Owl Capital. The silicon doesn’t care who was right first.

Zero times infinity isn’t a number. Mathematicians call it an indeterminate form: the product depends entirely on how fast each side is moving. Right now the price is heading to zero, the volume is heading to infinity, and the product is a bill. It’s being paid by Nvidia’s shareholders, Oracle’s bondholders, the Gulf sovereign funds, the pension plans behind the private credit, and everyone whose laptop cost 17% more this year.

It comes due for them, not for you. And when it does, the zero gets a little closer.


References: