{"id":"2093040177711329673","url":"https://x.com/fi56622380/status/2093040177711329673","text":"https://x.com/i/article/2092862654117498880","author":{"name":"fin","username":"fi56622380","avatarUrl":"https://pbs.twimg.com/profile_images/1617438471773360129/PuNEnXyH_200x200.jpg"},"createdAt":"Thu Aug 27 18:17:19 +0000 2026","engagement":{"replies":31,"retweets":243,"likes":1522,"views":480185},"article":{"title":"AI Semiconductor Endgame 2026 (III) ","previewText":"The next ARR growth engine after coding, the AI infra bubble, the Economics of open vs. closed source model\nWhy is the AI semi buildout longer than the internet buildout?\nAfter coding, where does the","coverImageUrl":"https://pbs.twimg.com/media/HQvlJaobgAAUv0o.png","content":"The next ARR growth engine after coding, the AI infra bubble, the Economics of open vs. closed source model\n\nWhy is the AI semi buildout longer than the internet buildout?\n\nAfter coding, where does the next leg of ARR growth come from?\n\nWill AI infra be overbuilt? Is capex a bubble? Does the compute math pencil out?\n\nWhat does the open- vs. closed-source landscape look like? How will the impact on closed-source models evolve?\n\nThis piece is an attempt to reason through the logic behind these questions\n\n—--------------------------------\n\n## Why is AI's infra buildout longer than the internet's?\n\nHardware/infrastructure revenue growth in the internet revolution had exponential growth along only one dimension of penetration: user count\n\nBecause everyone was on a subscription — $50 a month for internet access, and there's nowhere to upsell from there. It's a flat-rate service. Looking at ten times as many web pages doesn't make you pay ten times as much. A flat-rate service has no second dimension of exponential growth; once it saturates, it saturates\n\nMobile internet infrastructure — the edge device, the smartphone — runs on the same logic. Edge devices can't carry a high ASP, because amortizing cost requires mass production, and to get mass scale the ASP can't be high, and it's hard to raise. So edge devices were destined to stop compounding once penetration reached a certain level, because the ceiling is determined by a single dimension.\n\nThis is also why hardware demand in the internet revolution inflected and slowed roughly after penetration passed the midpoint of the sigmoid. When the internet bubble burst in March 2000, US internet penetration was around 50%. The fastest stretch of any technology revolution is the move from 10% adoption to 50%.\n\n—--------\n\nAI hardware demand, by contrast, comes from the exponential growth of tokens. The reason this round of AI is so hardware-hungry is fundamentally that the ceiling is determined by two dimensions: exponential growth in user count/penetration, and exponential growth in usage per person\n\nThese two dimensions of exponential growth multiply, and both convert into a single unit of account: tokens — which raises the infrastructure ceiling by far, far more\n\nWhere does each of the two dimensions stand in our era?\n\nThe first dimension, user penetration, crossed 50% a few months ago — meaning the derivative of the sigmoid is about to inflect\n\n![](https://pbs.twimg.com/media/HQtYhwXawAAFI-p.png)\n\nThe second dimension, tokens per person, is still early in 2026. The median AI spend inside US enterprises is currently $12 per person per month. Long term, that number could plausibly reach at least 10% of a white-collar salary — roughly $1,000 a month. There are nearly two orders of magnitude in between.\n\n![](https://pbs.twimg.com/media/HQtYarZbAAAIBpf.jpg)\n\nToday's growth is mainly the second dimension of penetration, the exponential growth of average tokens per person. The two sigmoid curves have run a relay\n\nThis should be historically rare: compute/infrastructure demand created by a relay of sigmoid curves across two dimensions, and with it a semiconductor supercycle\n\nBut all exponential growth is an illusion before the ceiling arrives, and this round of AI buildout is no exception — the illusion just lasts much longer, because there's an extra dimension of exponential growth taking the baton\n\n—--------------------------------\n\nThe other reason AI's buildout demand runs higher than the internet's was mentioned in the semiconductor year-end piece I wrote last year: as noted, both of these penetration dimensions convert into a single unit of account, tokens. Unlike the internet era — where the network and compute cost of each request had extremely low marginal cost, scaling worked beautifully, the marginal cost of copying and distributing software was close to zero, and almost no extra hardware was consumed.\n\nIn the token era, because of the high cost of inference, the marginal cost of a user's access is an order of magnitude higher than in the internet era. The business model has shifted partly from \"traffic × conversion rate\" to \"gross margin per token,\" which is why infrastructure demand in the AI era is far higher than in the internet era\n\n—--------------------------------\n\nThe first dimension of penetration diffused so fast that it has already inflected, which is why, when people discuss AI's prospects now, penetration is no longer the focus. The discussion is much more about how tokens per person will grow across every industry.\n\n![](https://pbs.twimg.com/media/HQtYtBRa4AATnNm.png)\n\n—--------------------------------\n\n## After coding agents proved their value with ARR ramping fast, growth started slowing in June–July, and the market began to worry: once coding agents have broadly penetrated the developer population, what scenario takes the baton on token/ARR growth? Is it AI for science? Or RSI?\n\nI think it's still coding — coding in the broad sense — and there's still a lot of room\n\nCoding is the unstoppable trend of the world going digital. Everything can be coding, and it doesn't have to be explicit coding. In practice, far too many tasks are quietly being converted into coding in the background, and when the work is logged, they still show up as coding.\n\nRight now, broadly defined coding tasks account for roughly 60–70% of all ARR revenue. If you count only the narrow developer definition it's maybe 40%; the other 20–30% is likely tasks from other industries converted into a coding-agent execution structure.\n\nTake the most common example: editing files, Excel and PowerPoint — in the end they're all completed by spinning up a Linux sandbox in the background and doing it via code. Finance, CRM operations, legal document organization are the same. The only difference is which tools and verifiers get called. All of it is converted into observable -> executable -> verifiable closed-loop coding tasks.\n\nSemiconductor architecture exploration converts into a SystemC model simulation problem, many steps in molecular structure analysis and simulation in biopharma can be done with code, and model-building inside an investment bank's research report can be turned into a coding problem too.\n\nThe narrow developer coding task is simply the first large task category to have gotten the high-frequency observe/analyze -> execute -> verify loop running end to end, which is why coding consumes an order of magnitude more tokens than chat. There are far too many tasks that will inherit the same logic later — including open-ended tasks like deep research. In form, the next realistic source of growth is co-work/computer use\n\nWhat takes the baton from developer coding is not some other isolated mega-scenario. The fastest-growing thing right now is precisely \"non-programmers using coding infrastructure to complete non-coding product tasks\", coding is still the default execution kernel, but it isn't necessarily the product category the user sees\n\n![](https://pbs.twimg.com/media/HQtZRCab0AEBsqD.png)\n\nAs of June 2026, among OpenAI's enterprise customers, Codex already accounts for 64% of combined Codex + ChatGPT output tokens. Since February, Codex weekly active users have grown 108x in legal, 41x in sales and recruiting, 26x in marketing — and only 5x in engineering\n\nLooking at Anthropic's revenue mix by vertical, narrow developer/software-company coding is only 40%, financial services/insurance is over 20%, and legal, pharma/life sciences, and consumer/retail e-commerce are all very substantial, each on the order of 10%.\n\nThe logic of software eating the world, once LLMs lowered the barrier to coding by a hundred times, has even intensified into coding eating the world, or put another way, coding is the modality agents are most fluent in, and agents are using coding to spread their utilization into all knowledge work\n\nAgent coding is already past the feasibility and direction-finding stage. What's left is mostly engineering problems — long-running execution, cross-system operation, autonomy, reliability. No major challenges remain.\n\nThe rest is a matter of users adapting to the scenarios and changing their habits. The main bottleneck on penetration/utilization is that people's learning curves aren't that fast, users need time to migrate to agent/coding workflows; enterprise workflows are highly long-tail, data is scattered, internal tool integration is insufficient, and there's a lack of top-down KPI pressure. Model capability, organizational push, and tool integration jointly determine adoption speed.\n\nThe other bottleneck logic is Amdahl's law (the outcome is governed by the part you can't accelerate). Tokens only accelerate some of the steps — the clinical portion of drug development, for instance, can't be accelerated. If half the steps can't be sped up, then even if tokens accelerate the rest a thousandfold, the overall speedup is only 2x. So the more realistic path is to explore more directions in parallel, moving the bottleneck from decision-making to verification.\n\nAnthropic's revenue slowdown in June–July had three main causes: not enough compute; an anti-distillation crackdown that shut down a lot of accounts; and OpenAI's Codex 5.6 taking share. It was not broad agent-coding demand hitting a ceiling.\n\nFrom the broad agent-coding angle, I don't currently see a hard ceiling on large-model/agent demand. We're still in the phase where the more you use it, the more efficiency you gain — the marginal positive return is still large, and the main bottleneck is how much budget revenue/profit can support. Willingness and use cases aren't the problem.\n\n----------------------------------------------------------------\n\n## Reasoning from the demand side, which matters most, there's nothing wrong with the outlook for token growth rates. So reasoning from the side that funds the buildout: where is the ceiling on capex, can it be sustained? Is there a bubble? Debt plus FCF turning negative — does the math work?\n\nFrom the funder's side, the biggest difference from the internet is that in the internet era the parties funding the buildout and the parties benefiting from it were not the same. In the AI era, the two overlap. The largest funders of the buildout, Anthropic and OpenAI, are also the largest beneficiaries of this AI revolution.\n\nSo the most direct way to run the math is from OpenAI's and Anthropic's point of view\n\n1. The unit economics work\n\nAs of end-June, Anthropic generated about $62B of ARR on roughly 2GW of effective compute. If you put the build cost of 1GW of datacenter at $50B and amortize it over six years, the annual cost is about $10B. So Anthropic's inference compute revenue of ~$30B/GW divided by the annualized compute cost is roughly 3x. Inference ROIC is around 60% as well\n\nIf you count inference only, the unit economics look even better — revenue per GW is close to $50B\n\n2. As long as that unit-economics standard holds, the datacenter compute GW plans of Anthropic + OpenAI pencil out, so long as ARR growth stays in step with compute growth\n\nIn July the two of them had roughly 6–7GW combined; Anthropic $65B, OpenAI $45B\n\n![](https://pbs.twimg.com/media/HQtZxwJbAAEciBx.jpg)\n\nBy year-end the two will bring roughly 11GW online in total; ARR only needs to grow another ~60% over the half to keep ARR and online compute growing in step (keep in mind that in 1H26 A did 7x and O did a bit over 2x) — that is, 10% per month in the second half\n\nBy end-2027, the two companies' planned compute totals roughly 20–24GW. In fact token ARR only needs to grow 1–1.5x over that year, and the compute math pencils out at least on unit economics. Compute and ARR growth stay in step, and current visibility supports 2027E capex without much problem\n\n2028, where there's no visibility, can only be guessed at: if 2028E ARR grows 80% in the year, that justifies a combined 45GW of compute for O+A. If ARR grows another 100% across 2029 + 2030 (a CAGR of just 40%), then O + A's roughly 100GW of total planned compute by end-2030 is entirely reasonable\n\nAt the very least, for players like O and A that have clearly found a structurally growing demand lane, the ROIC math works. The 2027 and even 2028 compute plans are not, for now, a bubble\n\nIs taking on debt and going FCF-negative to build AI infrastructure a gamble? Yes\n\nBut that kind of gamble is rational in front of a business with ROIC this good and a high probability of persistence. The risk of missing the move is meaningfully larger than the risk of an FOMO bubble. Even if financing costs rise to 8% or higher, at least through 2027–2028 the return on compute can still justify the borrowing risk.\n\n—------------------------\n\n- The compute math works for O and A. What about the rest of the compute? Does the return pencil out?\n\nThis part is very hard to quantify; you can only do it roughly and qualitatively, because a lot of compute generates no direct revenue and is booked as R&D cost. Meta, for example, plans to add 5GW next year for MSL model training and search/ads/recommendation, and Google's Gemini training plus search/ads/recommendation is the same — none of it produces revenue directly, but these are still strategically reasonable compute needs.\n\nHere's something easy to overlook: industry-wide, a very large portion of AI/ML compute today isn't selling tokens at all — it's still traditional search/advertising/recommendation, Meta being the classic case, and most of Google's compute works this way too\n\nSo when you look at the astronomical compute numbers for 2027–2028, traditional workloads are an important and easily overlooked component. Inside the Silicon Valley companies along the supply chain, the feel is the exact opposite of the market's: internally they are desperately short of compute, while the market is still questioning whether the compute math shows any revenue, and worrying about oversupply.\n\nBut once you count in all of this R&D that generates no direct revenue yet still matters, plus the large growth in traditional AI/ML compute, the portion of compute that doesn't pencil out is quite limited.\n\nOn top of that, O + A's share of compute keeps concentrating: for net-new compute GW across 2027, the power bottleneck means the US can only energize about 20–25GW, and O + A may take 10GW of that — perhaps close to half of the net-new. In 2025 that ratio was lower: O+A was only about 25% of net-new compute, and a large share of compute GW wasn't LLM at all. As time goes on, O + A's share of compute GW keeps rising, which means the portion of total compute whose math works keeps rising too.\n\nAnother easily overlooked piece of logic: because coding agents took off, the iteration speed of the software industry and the ML industry suddenly accelerated, the digitization of the whole world accelerated, and the traditional hardware infrastructure demand that comes with it is accelerating too\n\nTake CPUs as an example:\n\ntoken = decision/intelligence\n\nCPU = execution/verification\n\nIn a digital world where everything is coding, as decision-making capability explodes, demand for execution/verification grows in lockstep by the same order of magnitude\n\nAnd the faster iteration in the traditional ML industry follows CPU-like logic too: as ML results improve and use cases multiply, that also drives more GPU/ASIC hardware demand\n\n—------------------------\n\n## Will AI infrastructure produce a bubble? Yes — and it will definitely end in overbuild\n\nInference is simply too profitable right now, so overbuild is an inevitable outcome, because everyone wants to grab a piece of the inference market, every player is in the arms race, nobody can stomach losing share and failing to ride the wave of high-speed growth, and no one is going to back down\n\nClosed-source inference gross margins clear 70% easily (even after the CSP takes its cut), and token factories hosting open-source models and selling tokens run 60%+ gross margins as well (DeepSeek's GM is beyond 80%), GPU prices keep rising, rent for 1GW is heading toward $20B, and even just selling short-term GPU cluster contracts carries 60%+ gross margin. This AI infrastructure market is so profitable and so tempting that anyone with distribution and technology is piling in as a neocloud, and because there's generally a one- to two-year gap between build plans and delivery, overbuild is certain to happen — it's only a question of when\n\n\"A question of when\" means that at current demand and current compute plans, overbuild isn't visible for at least two years. Any step-function improvement in models (reasoning -> agent, for instance), or any innovation in the application paradigm that sustains AI spending growth, keeps extending the overbuild timeline further out\n\nWhen those extensions can no longer keep up with the pace of the buildout, overbuild will slowly start to surface\n\nBut this is actually better for the leading model companies and the application layer, because overbuild pushes datacenter gross margins down, and more of the inference profit flows to the top two\n\nSpaceX is still small in scale, but in this brutal market for building new compute it has the strongest execution and the lowest cost, possibly the strongest competitiveness, and the biggest ambition and plans. It's bound to become one of the most important contributors to overbuild.\n\n—-----------------------------------------------------\n\n## Will the rise of open-source models affect closed-source ARR? What does the landscape look like today?\n\nWill the gap between open and closed models narrow?\n\nWhether open source can sustain its momentum also depends on open-model capability. My view is that the gap between open and closed models will hold at around 6 months (Mythos came out in February and got regulated). Distillation is mainly about cost reduction\n\nThree moats in model progress: compute scale, data scale, self-iteration\n\nOpenAI and Anthropic together spent over $10B on data this year, but Chinese labs are now spending billions of dollars on data too. The data-scale moat has narrowed, but it's still there\n\nOn compute, between building datacenters offshore and transshipment through various countries, the GPU ban has largely been worked around. Once domestic Chinese accelerator capacity ramps, the compute moat will keep narrowing too. For now, the order-of-magnitude gap in compute still holds\n\nRight now, progress on the next generation of models depends heavily on iteration off the previous generation (self-iteration). The stronger the last generation, the stronger the next — that gap is not easy to close\n\n—-----------------------------------------------------------\n\nAs for open source's impact on closed-source ARR, there certainly is one, but it may be smaller than the market expects, at least within a year the hit isn't large. Let me try a rough quantification here\n\nAs for open-source token share on OpenRouter and Vercel, going from 28% to 62%, that only reflects the shifting picture among small/micro businesses and multi-model users on routers; the vast majority of closed-source tokens simply aren't visible there, so the sample doesn't count for much\n\nOn ARR as of end-July, Anthropic is roughly $72–75B and OpenAI is roughly $45B\n\nAs of July, all open-source revenue globally adds up to about $10B; North America alone is only $4B, $5B at most. On an ARR basis, North American open-source token factories + AWS Bedrock open source + Azure open source is only <$4B of ARR, against $120B closed source — a ratio of roughly 3%\n\nBreaking down China's $3.5B of open source: roughly Zhipu $1B, Alibaba $1.2B, DeepSeek $500M, Kimi $300M, MiniMax $200M. Breaking down the US: Fireworks $1B, Together $1.2B (open-source inference is only $0.5B of it), Baseten $0.5B. Add in other regions and you barely get to $10B\n\nNorth American token-factory open-source inference grew only about 3x in 1H26 (Fireworks, for example, went from 10T+/$300M ARR to 40T/$1B ARR). Overall, compared with the growth rate of the closed-source inference market, there's no meaningful advantage\n\nAs for enterprises deploying private open-source models for their own use, that path doesn't work from a cost standpoint (data security is a separate consideration). If the open-source model isn't optimized well, the cost can even exceed closed source — you're better off handing it to a token factory\n\nUS open-source compute expansion is still limited by compute, so it can't scale up, and its expansion speed has no advantage over closed source. Fireworks, Baseten, and Together have all raised on the order of $1B, and at current compute rental prices of $20B/GW/year, that rents less than 0.5GW for a year. Their compute growth isn't actually faster than closed source. To expand, they have to get customers to bring their own compute on Azure/AWS\n\nChina's open-source expansion is faster, but its token factories are equally compute-constrained — they're nowhere near meeting China's own demand, so for now they can't go abroad and sell to the US at scale. China has roughly 30 token factories of various sizes, already a cutthroat red ocean, all severely short of compute: whoever has cards has revenue\n\nAnother risk for Chinese tokens going abroad is trust and policy risk. US enterprises may not dare to use them, and China may not want its own models used for free with no return. But as routers like OpenRouter and Vercel keep developing, even a ban on exporting tokens won't be able to stop tokens from going abroad, since the profit on selling open-source tokens to the US is simply too high\n\nThe problem with US open-source token factories is that compute is limited for now and financing scale is limited for now, so expansion speed can't meet demand. Even if policy opened up and Chinese open-source tokens could be sold abroad, there isn't enough compute in the short term for them to get out. This is also why Nvidia is working so hard to support open source, even stepping in directly to accelerate the process (it acquired Hugging Face today)\n\n—-----------\n\n![](https://pbs.twimg.com/media/HQtaT1-a8AArTv_.jpg)\n\nSeparately, the closed-source defensive playbook is quite strong — 5.6 Luna cutting price by 80%, for example, makes it very price competitive\n\nWe can compare it to history: this is exactly the same defensive strategy Qualcomm used against MediaTek for years in the mobile internet era:\n\nSol is the equivalent of the Snapdragon 8 series: the flagship holds the technology and brand ceiling, and the volume and share aren't that large\n\nLuna is the equivalent of the Snapdragon 4 series: proactively fill the low-price band, leave no profit pool for low-cost competitors, and most of the volume comes from here\n\nJust like tokens split into different tiers, MediaTek's chips were in fact far beyond \"good enough\" for a phone — never mind everyday use, even in high-load gaming scenarios MediaTek wasn't far behind Qualcomm a decade ago\n\nThe Qualcomm–MediaTek offense-defense war ran for 15 years, and with this strategy Qualcomm still defended its market share, held up gross margins on high-end chips pretty well, and remained the first choice for Android flagships\n\nSo looking at history, open source's impact on closed source is real, but it won't change the course of history\n\n—-------------------\n\nAnother open-source-adjacent question: will the end of the tokenmaxxing wave affect ARR growth?\n\nI think tokenmaxxing was just a short-term thing companies did for internal adoption purposes. Some large US firms do currently cap employees' daily usage, but the cap is very high — roughly $10k–20k a month — and over 95% of employees don't come anywhere close to using it up. The median is in fact very low, so there's still plenty of room to grow. Token usage is still in the phase where the marginal positive return is large\n\n—--------------------------------\n\nOverall\n\nThe reason the outlook for AI infrastructure is so murky, and the bubble debate more heated than in the internet buildout era, is mainly that compute spending has certain visibility while demand visibility is low, which creates enormous uncertainty, and in between sits a layer of financial risk from the timing mismatch between revenue and spending — a buildout financial risk that, as every party keeps raising its bet, has escalated to the point where it can even affect Treasury issuance\n\nBut I hold to the view that even on a base case with no new paradigm at all, the buildout through at least 2027–2028 is not a bubble\n\nCompute is intelligence, and intelligence is enormous revenue.\n\nThe speed of the relay between AI token demand's two exponential dimensions is without precedent in history; the demand from digitizing everything through broad coding has only finished its first leg; open source's impact on profit distribution is bound to happen, but it won't change the overall course\n\nFrom the supply-chain angle and from the market-macro angle, the risk and the sense of how good the cycle is look opposite, but the upheaval both sides feel is equally violent: those up close see not enough, those far away see too much\n\nCompute has become judgment/cognition. How to project demand, how to calculate the ceiling — there is no historical reference at all. What we're really arguing about was never whether the number is right, but what unit of time to look at it in\n\nAnd that fact itself is precisely what shows how big this is: only a true phase transition makes all the old historical references fail at once\n\nThe visibility of spending creates today's fear, while the invisibility of demand is exactly where this era's biggest opportunity is hidden\n\nNot seeing the path clearly doesn't necessarily mean seeing the direction wrong\n\nEvery S-curve exponential eventually reaches its end.\n\nBut civilization has never stopped at the end of any single curve.\n\n—--------------------------------------\n\nWhat is open source's impact on semiconductors?\n\nWho wins the semiconductor capex civil war? And if memory/storage eats too much of the profit, what happens?\n\nThis piece is already too long, so those two topics roll over to the next one"}}