Essays Tim Davis Essays Tim Davis

The Token Curve

Heterogeneous compute will materially reprice the long-contracted GPU market.

Financing terms will set the cost floor for every token served, and I believe heterogeneous compute will materially reprice the long-contracted GPU market.

Previously, I wrote about GPU financing, debt markets, and how software fungibility can make compute liquidity broader and increasingly cross-vendor over time. I wrote that essay and this one at the same time, then split them so each could make a cleaner point - and kept this one shorter. Here, I focus on what I call “The Token Curve” - how heterogeneous compute is going to reshape and reprice the market for long-term GPU commitments.

The core argument is that portability does not create heterogeneous compute - rather, it turns the heterogeneous capacity that already exists into supply the market can actually use. Once enough of that capacity is production-ready and available at meaningful scale, I believe it will reset the marginal cost of inference and reprice fixed GPU commitments.

In this essay, I consider “demand” as having three different measurements, and they’re growing at different rates:

  1. Usage volume - tokens actually served - is compounding at something absurd, on one of the fastest adoption curves the industry has recorded.

  2. Inference revenue - the product of rapidly rising volume and rapidly falling realized prices, so it can grow much more slowly than token volume.

  3. Infrastructure investment - capex - is neither of these; it’s a forward capital commitment against forecasts of both.

Visually you can think of this looking something like this:

FIG. 0   This is just an illustration of three ways to read demand, as illustrative shapes rather than data: usage volume compounding fastest; inference revenue as volume x falling prices; infrastructure investment in step functions committed ahead of both. A more measured versions appears in FIG 01 and note 1.

When I hear people say “demand isn’t slowing,” I think they’re usually pointing at capex, which in my view is the least direct and most forward-looking proxy of the three - it measures capital committed against expected demand rather than consumption itself, even when much of that spending is already supported by customer contracts. Instead, I wanted to focus on the gap between the second and third measurements - the distance between the inference revenue actually arriving and the capital being committed against it - a gap that I noticed the Bank for International Settlements (and I’m sure many others) has begun measuring versions of this directly (which my first essay walks through on the credit side).[1] That financing gap sets up the token curve, as I use it: the relationship between two curves this essay will separate - the market price of fixed-capability inference and the portable break-even cost of producing it - and the basis between them - asking, what will a token of a given capability cost to produce next quarter, next year, and the year after? Think of it as the forward-looking cost of producing the same unit of intelligence.

One practical thing I’ve learned is that financing terms do not set the market price of a token - competition does. By this I mean that an operator’s required cost floor can be mapped in a simple equation: required token cost = (capital recovery + power + operations) ÷ effective tokens produced. With token cost defined this way, we can watch how each variable moves the equation over time.

Interestingly, this also tells us that falling unit prices can coexist with massive revenue when volume scales up - the risk being priced becomes the unhedged mismatch between a fixed cost line and a repricing revenue line. To make it simple, I just use the standard industry language - dollars per million tokens - and then try to understand what a “collapse” looks like. Using public data, a clear example is that Alphabet’s capex intensity per token served has fallen almost ninety-fold in two years, and the only way that happens while capex nearly quadruples is because the underlying demand denominator grew 330x in the same period.

FIG. 01   Scene-setting in two panels on a log-scale: usage volume above; below, two directional indicators - Alphabet capex intensity (a ratio) and an illustrative constant-capability envelope, with mid-2026 open-weight parity releases ticked - not comparable unit costs. Methodology in note 1; envelope construction and parity dates in note 4. Lower-panel series share the dollar axis for direction, not level.

The math behind the above chart is pretty simple, using a log scale - the demand line is Sundar’s own numbers from three successive I/O keynotes: Google served 9.7 trillion tokens a month in May 2024, roughly 480 trillion in May 2025, and over 3.2 quadrillion in May 2026. The capex line comes from Alphabet’s filings: $52.5 billion in 2024, $91.4 billion in 2025, and a 2026 guidance range of $195 to $205 billion, of which I use the $200 billion midpoint. Divide each year’s capex by that year’s May token run-rate annualized - $52.5 billion over roughly 116 trillion tokens, $91.4 billion over 5,760 trillion, $200 billion over 38,400 trillion - and you get roughly $451, $16, and $5 of capex per million tokens served. The market-price line is what I’m calling the “constant-capability envelope”: an illustrative price path for a fixed reference task at a defined quality threshold - a constant-capability API-price envelope rather than a measured “task-cost” index. An actual basis contract would translate both revenue and serving cost into dollars per standardized task using a fixed input, output and caching profile. I anchor it at GPT-4-class output, roughly $60 per million tokens in March 2023, falling about 10x per year since - the same slope a16z documents for GPT-3-class output from 2021, with task-level rates varying 9x to 900x per Epoch.

Every later model example in this essay is an observation within this envelope and I will keep referring back to it throughout. Some caveats in the data I have access to are that capex is an intensity ratio, not a unit cost, reflecting that it is spent ahead of usage. It also includes both training and infrastructure that never actually serves a token and I would also note that Google’s token count spans every surface it runs - so not just just paid inference as another qualifier. The two series on the chart are therefore not directly comparable unit costs - my point is that the figure demonstrates the demand denominator and the direction of market pricing, not Google’s inference economics. I don’t believe any of these caveats change the underlying thesis however. I want to also make it clear that there are two different curves in play - the market curve and the portable cost curve:

  1. The market token curve is the price of completing a standardized task, regardless of which model does it - open-model competition moves this the most.

  2. The portable cost curve is the break-even cost of running a fixed checkpoint at a defined fidelity and service level across whatever qualified hardware is cheapest - and it’s my view that silicon portability moves that one.

We can test whether the same model behaves consistently across different hardware, but comparing different models requires judging whether they can perform the same task at the same quality. For any provider, the Token Curve is the gap between what the market will pay and what it costs them to produce that work.

If we combine the learnings around how the debt behind the AI buildout is actually priced, with the uncertain signals that lenders are providing in underwriting the silicon - then this essay is meant to map that uncertainty back on the actual dollar per token value over time, which is fundamentally what everyone is trying to understand long term.

The token sandwich

Let’s try to understand the math behind the price of a token.

To start, consider that a datacenter operator’s GPU-hour rate has to recover its capital inside whatever window the lender allows, plus the spread the lender charges: that is, when loans fully amortize in under five years, that entire recovery is compressed into 36 to 60 months of rentals. Not every company borrows money to buy its GPUs, and hyperscalers often pay for them with their own cash, but the economics are still similar: the company decides how many years the GPUs have to earn back their cost and what return the investment must produce. For a borrower, a lender sets that window, but for a cash-funded deployment, the company’s own financial modeling has to set it. Either way, the reality is that the hardware has a limited period in which it must pay for itself. We know this from the prior essay where CoreWeave takes roughly 26 cents of interest expense per revenue dollar - a company-level ratio rather than a literal per-token charge, but financing expense is ultimately recovered through fleet economics. Thats to say that it feeds the cost floor under every hour its GPUs rent, and every token served.[2] In CoreWeaves case, this burden persisted into the second quarter: $640 million of net interest on $2.575 billion of revenue - roughly 25 cents per dollar - alongside $1.393 billion of depreciation and amortization, which is most of the distance between its 59% adjusted EBITDA margin and its 5% adjusted operating margin (note 2). So its clear that financing and capital recovery sit underneath literally every GPU-hour.

But the real question is who is buying those hours? We can’t think of it as purely GPU hours, as the currency of AI is tokens and increasingly everyone wants to sell tokens, not GPU hours because of the higher margin. So in this regard, downstream of the datacenter operators sits a layer of inference providers - many of whom remain economically and operationally anchored to CUDA and NVIDIA capacity, even when they have begun adding AMD or multi-cloud support - and a common operating model is to procure compute on multi-year reserved compute cycles and sell it back to the world by the token.[3] For existing inference providers, the way to think about the trade they are making is that they are “long” a fixed capacity (despite being exposed to market repricing), and “long” their own efficiency roadmap.

Practically, that looks like locking in three years of Hopper or Blackwell at today’s reserved rate: cost of goods fixed on NVIDIA’s release schedule, while revenue is repricing at the market rate. The challenge here is that the market schedule is brutal - the price of constant-capability inference has fallen at something like 10x per year (on the assumptions I use) - my envelope anchors at GPT-4-class output from March 2023, and a16z documents the same slope for GPT-3-class output from 2021, a thousand-fold in three years - with task-level declines ranging from 9x to 900x annually - against the roughly 23% yearly decay in the GPU-hour itself.[4] So the argument doesn’t depend on some crazy steep and negative reading - the chart below is a 2023-2026 model I created with Claude. It holds a locked contract flat and divides it by the provider’s own efficiency roadmap - effective cost per task is the contracted GPU-hour divided by tasks per GPU-hour. As you can see, even 2x-a-year serving gains leave a locked contract indexed at 12.5 against a market band of 0.1 to 3.7, and repricing onto new contracts narrows but does not close the gap against a 3x market. The “sandwich” gap closes only when the efficiency roadmap exceeds the market curve, which is precisely what I'm trying to describe below.

FIG. 02   Two curves, one trap - a 2023-2026 backtest indexed to 100 (log scale): the illustrative market envelope as a 3x-10x scenario band, a locked contract divided by 1.2x and 2x annual efficiency, and new contracts at 2x efficiency. A locked book at 2x still indexes to 12.5 against a market band of 0.1-3.7. Sources: a16z LLMflation, Epoch AI (note 4); observed H100 reserved-rate path.

This pattern is holding at the current frontier - the mid-2026 wave of open-weight releases now benchmarks alongside the closed flagships at a fraction of the price, and capability parity now arrives months behind the frontier, with every locked reserved rate marked accordingly. Consider total cost per standardized task against a frontier open-weight model - we find that Kimi K3 is burning twice the median output tokens on Artificial Analysis’s index and still lands at $0.94 per task against Claude Opus 4.8’s $1.80. In the cited benchmark set, the open frontier is at or below closed pricing per task, not just per token, and token prices are falling faster than hardware prices because model efficiency compounds on top of hardware deflation.

I strongly believe this “token sandwich” is real - your costs are contractually fixed, and the market value of the work those GPUs produce can fall beneath them even while GPU-hour prices themselves remain elevated: your realized price per standardized task falls an order of magnitude faster than anything on the cost line. This makes the “sandwich” really, as in any wholesale-retail business: the margin concentrates in the long tail of smaller customers paying list prices, while the largest deployments - the volume that actually fills the pareto distribution of commitments - consistently negotiate toward cost. The duration-matched revenue is the thin-margin revenue while the fat-margin revenue is spiky, short-duration, and has no real commitment at all because it will just keep flipping between inference providers over time based on lowest cost.

So the obvious question is - how to eat the sandwich?

In my view, the obvious way out is to make the fleet operate even faster and smarter than the rate at which both curves decline - which is precisely the stated trajectory across the current inference market, where individual providers claim roughly “50% to 60%” gross margins on the strength of a proprietary inference stack - Fireworks’ disclosed ~50% blended with a stated 60% target being the “documented case”. Obviously inference is a real and massive ARR-growing business, but it is ops-alpha, and it compresses as open-source serving engines close the gap from below in a game of cat and mouse. In my view, increasingly the alpha is decomposing into three primary levers - quantization, speculation, disaggregation - each capable of order-one gains under the right workloads, and each commoditizing every time a new model ships because quantized checkpoints ship within days of a release, and the newest open models now ship their own speculation heads in the weights.[5] In my experience, and based on discussion with a lot of friends and enterprises in the industry, the “levers” that earn the margin are the same levers being given away, which means that alpha has to be re-earned every quarter. I do believe that “optimization alpha” is compressing really quickly - kernels, quantization recipes, speculation techniques, scheduling - because open-source engines reproduce them incredibly fast - while “platform alpha” is more durable: reliability, supply access, networking, storage, security, observability and the ability to run heterogeneous infrastructure as one system. CoreWeave attributes its premium pricing to exactly this, and books over $400 million of ARR from storage, CPU, networking and software (note 2).

Ultimately, my belief is that the more durable strategy over time is to own beta - a new structural position in the stack that has more asymmetry and changes the economic game. This is likely to come from some new hardware innovation, or from a completely new model architecture, or broader research optimization that hasn’t been invented yet. If you consider that even frontier open models are now profiling their own serving stacks and writing replacement kernels that run in production - this is a loop where the serving alpha is generated by the model being served and will have diminishing returns.

The reality of scaled production workloads is stranger still - the same model, on the same software, on the same GPUs, can run cleanly in one datacenter and break in another - the variation in quality across “neoclouds” is honestly incredible - and tiny timing differences in the network are enough to trip bugs in the code running on the same chip. Providers sometimes resolve this by pinning hardened deployments to the specific clusters where they were validated, and if a machine’s behavior depends on where it sits, it’s not really an interchangeable asset - it’s a component of one specific system, priced as if it were interchangeable.

The world becomes heterogeneous

All this said, my first essay’s point applies to buyers as much as GPU owners - fundamentally, a locked-in commitment simply looses value on someone else’s timeline. The S curve of technology innovation always yields enormous margins in the beginning because the gradient of change is so steep, but over time these diffuse throughout the market and competition yields much tighter returns.

I believe this backtest should make one consider the forward implications, and the five-year forward view makes the sandwich worse, because the commitments being signed today were priced in a single-vendor world, but the industry generally accepts that they mature into a heterogeneous one. Heterogeneity fundamentally has four claims - vendor (whose silicon), architectural (what kind of machine), phase (prefill, decode, cache), and generation-and-site (which vintage, in which building) - and the missing understanding in the market is that they compound. The reality is that announced capacity is not deployed capacity, and deployed capacity is not capacity available to third parties. I was curious of public announcements and I tried to map these commitments with dates (so don’t take this as a true market-share forecast). If you check the chart below, what it means is that the contracts locking in Hopper and Blackwell through 2029 and 2030 will be marked against a supply landscape that looks nothing like the one they were signed into. We already have obvious signs of this movement today - external TPU sales and a Blackstone-funded TPU cloud, AMD’s MI450 generation at gigawatt scale for OpenAI, Anthropic’s 3.5 gigawatts of TPU capacity, Cerebras at wafer scale, AWS custom silicon - led by Trainium - reaching a $20B annualized run rate, Qualcomm’s rack systems with high-bandwidth memory and NVIDIA’s Rubin cadence resetting its own price-performance every year regardless.

Here is the timeline for each commitment, its window and its confidence (from the announcements I could) find illustrated below[6]:

FIG. 03   Announced heterogeneous programs mapped to deployment windows - deployed, contracted/committed, product cadence, and announced/JV kept distinct, with the Trainium row splitting its deployed run rate from forward committed demand. Not a comparable supply curve; deployed is not merchant. Sources in note 6.

In December, NVIDIA itself paid roughly $20 billion to license Groq’s inference architecture and hire away its founding team - which is a significant tell - a $20 billion position on a rival inference architecture, which I read as a major hedge consistent with expecting a more heterogeneous inference market. AWS jumped in as well; as the largest cloud, they now pair their own Trainium silicon with Cerebras wafer-scale systems inside a single inference service - prefill on one vendor’s chip, decode on another’s - while its custom-silicon line runs at a $20 billion annualized pace against $225 billion of committed demand. When both the largest GPU incumbent, and the largest cloud, are engineering around a multi-architecture future - the direction should be increasingly clear to everyone. Interestingly, as GPUs grow more ASIC-shaped every generation (tensor memory, tile-scoped programming, phase-specialized SKUs like Rubin CPX) - many new entrants are trying to stay general enough to survive the next major architecture variant. The AWS-Cerebras pairing shows the heterogeneity is already happening but it’s software programmability that will turn this dual-architecture into a market-wide competitive unlock.

All this said, it’s interesting to reflect that any locked commitment today is always a bet against the future market price - and that price is about to change definition, from “next year’s NVIDIA chip” to “the cheapest capable silicon available from an expanding field.”[7] In this world, the floor resets downward every time the token leader changes - a lower envelope falls not because each curve is steeper, but because someone new keeps lowering it. The providers whose economics and deepest optimizations remain tied to fixed NVIDIA capacity are both the most exposed to that frontier and the least able to reach it, because a CUDA-shaped serving stack cannot readily route to the silicon that sets the new token floor without material re-engineering. If this world is true, then these folks become stranded twice over - they have supply contracted above the market, while also technically outside it from an engineering standpoint - a new entrant with no legacy commitments, and a portable stack, builds directly on the frontier and is so competitive on price that it can win a substantial share of the largest inference demand.

And heterogeneity is not only arriving between vendors - it is emerging inside the workload itself, as models add chain of thought, agents and continuous learning, each with a different computational shape - a cumulative “staircase” pattern that DeepSeek’s founder Liang Wenfeng describes in similar terms.[8] I made the same argument from two directions in What We Owe the Minds We Create and Scale or Surrender - the path forward is not a march toward one universal chip, but toward a widening portfolio of compute that must behave like one machine - and portability is what turns a collection of what we called at Google, “tech islands”, instead into one coherent and connected system.

We can already see the first version of this decomposition in inference itself - today, inference splits into two phases with opposite compute appetites → prefill, which is compute-bound, and decode, which is memory-bandwidth-bound - with the research record showing that serving them on different hardware is worth multiples, not percents - up to 7.4x more requests within latency targets in the DistServe work, 2.35x the throughput at the same cost in Microsoft’s Splitwise, and in production, Moonshot serves Kimi on a disaggregated, cache-centric fleet that handles 75% more requests, with gains up to 525% in long-context scenarios.[9] NVIDIA has now validated phase specialization in silicon - Rubin CPX is a prefill-specialized chip - GDDR7 instead of HBM, no NVLink - shipping at the end of this year paired with standard Rubin for decode in the same rack, and the company’s own framing at GTC was rack classes assigned to the phase each is cheapest to run - that confirms the phase split, not cross-vendor portability; the AWS-Cerebras pairing is the cross-architecture evidence. One possible strategic reading of the Groq license sharpens it further - SRAM-based LPUs for decode, GPUs for prefill - and on that reading, the $20 billion was NVIDIA buying the other half of a disaggregated future.

When the workload itself decomposes, the cost-optimal fleet is a portfolio of compute approached - mixed memory profiles, mixed compute densities, and mixed generations, because decode and cache tiers are exactly where older, cheaper silicon earns at the long tail of the market. The demand side supports that long tail too, as enterprises keep running the exact model version they tested and approved, often for years - it is common to find three-generation-old checkpoints still handling production batch jobs, because getting a thousand internal stakeholders is a huge pain in the ass. Old models executing on old chips is not a coincidence - proven production software that works is what keeps older hardware justified. The market is already pricing this: CoreWeave recently signed a contract for A100 capacity extending into 2029 at what it called an attractive price - a 2020 architecture (note 2). Jensen Huang even commented on this recently on X about the A100 lifetime of their hardware. However, it does not prove any particular A100 will earn for nine years - the purchase dates are undisclosed - but it gives the earning tail a concrete market example and re-emphasizes the point I made prior on software durability for chips.

Equally, a portfolio of compute ultimately needs a router - and the router’s scarcest resource is already clear in modern inference - KV-cache movement - state that must cross a network path whose bandwidth sits far below local HBM, often with extra staging and synchronization on the way - which is why the next hardware cycle is being designed around moving state rather than simply multiplying matrices together. Disaggregation turns portability into the architecture - so in my view, that world looks like even a single request, no longer having a single “best chip”.

So what happens?

The take-or-pay contracts that make GPU debt investment grade are clearly a fixed cost of goods, mapped against rapidly depreciating token prices. If token deflation outruns a platform’s optimization alpha, the margin squeeze arrives against the credit structure - CoreWeave’s DDTL 3.0 amendment showed the channel exists, as I highlighted in my prior essay. That is to say, customer delivery timing alone reached a GPU-backed covenant, requiring an amendment on one of its multi-billion-dollar facilities. CoreWeave’s financing now shows both sides of this transition. DDTL 3.0 showed that customer-delivery delays can reach an asset-backed covenant; DDTL 5.5 shows the reverse - lenders accepted a roughly five-year facility against customer contracts averaging about three years, leaving the debt dependent on renewal or re-leasing after the first contracts expire (note 2). That is early evidence of capital markets underwriting part of the infrastructure’s earning life beyond its first customer agreement - though it is still confidence in NVIDIA infrastructure on CoreWeave’s platform, not yet in a cross-vendor market for interchangeable compute.

If you follow the integration logic in both directions - the industry’s current shape starts to make sense. From below - in a market where the owner’s margin is often the most durable, every inference business is under pressure to become a miniature AWS or GCP - own the silicon, capture the infrastructure spread, and then sell the software on top - and the convergence is already visible in both directions.[10] CoreWeave is already living this: its managed-inference booked ARR went from roughly $1 million to more than $100 million within months, targeting at least $250 million by year-end - and management explicitly frames it as a way to put GPUs coming off contract back to work (note 2). CoreWeave is not merely exposed to the token sandwich; it is showing how an infrastructure owner can escape it by owning more layers of the stack. From above, it’s a fair bit scarier, as the model layer that consolidates to two or three labs earning software margins on inference can, over time, trend toward monopsony power over everything beneath it - power, data centers, silicon - and the biggest lab-silicon deals in FIG 03 are early illustrations of that.

So a world where every token seller must also become a data center owner is a world with a thin and incomplete market for compute - the difference between a concentrated model layer and a competitive one is not just margin - it is credit, as two or three natural buyers of the world’s compute is the A380 outcome (a reference in my prior essay to the aviation industry) applied to the entire fleet - collateral with a single exit. Fewer natural buyers mean thinner remarketing depth and lower recovery confidence today, while a competitive model layer - open weights at the frontier, vertically integrated entrants who treat the model as a cost center - is what “many exit buyers” looks like for compute collateral.

I learn that aviation did this in basically reverse - 2.4% of the world’s fleet was leased in 1980, more than half is today, as the asset became financeable, ownership unbundled toward lessors - financeability contributed to the shift rather than causing it alone - and operating an airline stopped requiring a balance sheet full of jets.[11] This is consistent with the incumbent vendor’s incentive to cheer for open source - in late July a 25-company letter, “Open Weights and American AI Leadership,” signed by NVIDIA, Meta, Microsoft, IBM and much of the American stack[12] - and why the two de-cornering forces travel together - open weights keep the model layer contestable, portability keeps the compute layer reachable. Within hours of a frontier open-weight drop, a dozen providers race each other’s tokens-per-second in public - the demand-side reflection of “many exit buyers,” running live every few weeks. A contestable model layer distributes bargaining power and demand more broadly through the stack - and broader demand is what could deepen the collateral confidence lenders need. The letter makes the competition, sovereignty, and security case for an open model layer; the credit case is the one it leaves clearly unaddressed. Vertical integration is best read as a response to missing markets rather than proof that owning everything is superior - companies own every layer of the stack because there is no reliable market that lets them buy the pieces instead.

Lastly, and I’ll make this same claim as I made in my prior essay - portability lowers the token-cost frontier and increases the pressure on market prices. If you compress the financing spread by the aviation-calibrated 100 to 250 basis points - my approximate numbers from the first essay - the hourly compute rate falls. This means you can stretch amortization across the asset’s real earning lifecycle, instead of a five-year forced period, and the hourly rate falls again. If you can lift utilization and arbitrage workloads onto whatever adequate silicon is cheapest - I predict it will fall a third time.

Every term of “the sandwich” improves for the buyer even as the curve gets steeper because portability converts a fixed bet into a “routing one” - it matters a lot that commitments are written against capacity “classes” instead of chip SKUs, and that a spot market exists that is fungible enough that the new compute futures become a hedge. Fundamentally, lenders finance token streams against a forward-looking curve - similar to the way power projects borrow against megawatt-hours - versus chips financed by credit. Ultimately, a token curve needs some type of grade, and tokens are not yet fungible across providers - so this means that quality drifts with quantization choices, speculators tuned to each provider’s traffic, and latency classes that differ. The financeable unit is not “a million tokens” - it is reference-model-equivalent inference at a defined quality and service level. The logit-fidelity tolerance, latency class, context length, throughput and uptime, geography and compliance - the way a real contract is specified. This first considers the portable cost curve as a fixed reference checkpoint and then secondly, considers task-capability benchmarks across unrelated models. Any mature market likely needs both, and so we need a better path to "certifying" this in the future.

The closing argument

Here’s my overall conclusion balanced against different arguments I’ve heard consistently from friends, and folks I’ve spoken too across the AI industry, that I thought I would address.

Portability is not free

The first argument comes from folks who are deep in the trenches of optimizing production workloads on LLMs - they consistently tell me that “portability is not free”. Every model-silicon pairing costs real bring-up and enablement time - quantization, a speculator trained from the base model’s own hidden states, parallelism and disaggregation tuned for interconnect, and then weeks of production hardening as live traffic finds what Artificial Analysis raw benchmarks miss for real customer workloads. Day-zero support means a token comes out - but the reality is that “production-ready inference” arrives much later, and it arrives per target. Multiply that by an expanding field and “cheapest adequate silicon” starts to look like a lot of unpaid engineering. I agree with that, and I’d note that cost is very high because a duplicative “fixed component” plus a smaller recurring per-target one, paid today per provider, per model, per chip - is exactly why its such a huge friction wall. Paid at a portability layer - amortized across operators and deployments, with a much smaller incremental cost per new target - the same cost is fundamentally a burden. While a portability layer adds its own abstraction, validation and coordination costs - the economic case I believe we are headed for is the amortized savings and routing optionality exceed them, and at fleet scale I believe they clearly do. The objection is really an argument about who bears the fixed cost of heterogeneity, not about whether the frontier exists.

FIG. 04   The enablement stack per model-silicon pairing, and cumulative enablement cost against the number of silicon targets - captive paid per provider per pairing vs a portability layer amortized with low incremental cost per target. Axes are relative engineering effort; illustrative, not a forecast (note 5).

GPU rates are spiking higher

The second argument I have both discussed extensively, and read about, is that the inversion of GPU prices is evidence that token prices are just going to get higher. Gavin Baker argued in an X post in late July 2026 that because spot GPU rental rates are running at least twice contracted rates, it means the hyperscalers are under-earning. His conclusion was that as contracts roll off and reprice higher, operating cash flow accelerates enough to fund the buildout, credit spreads compress, and the financing worry dissolves.[13] Indeed, if the world’s token appetite continues to explode and older GPU fleets rebook near original pricing - both claims my first essay makes - then locked commitments aren’t stranded, they are actually the only way to get capacity at all.

CoreWeave’s second-quarter results are the strongest evidence yet for the demand-heavy branch of this argument: near-term capacity effectively spoken for, prior-generation pricing at or above year-ago levels, a roughly 25% price increase across SKUs in July, and contribution margins on newly signed contracts expected to run five to ten points above prior quarters (note 2). That does not mean fixed-capability inference prices have stopped falling - it means GPU-hour scarcity and the market price of standardized inference can move in opposite directions, and may keep doing so while power and single-vendor supply remain the binding constraints.

The challenge to this argument is history - on the assumptions I’ve used here - constant-capability prices fell at roughly tenfold a year straight through the worst compute scarcity on record, and that argument survives the 3x and 5x cases too, because the deflation is driven by model efficiency and open-weight competition, not by hardware gluts. However, strong demand solves the “wrong half” of the problem - while it keeps a reversed compute fleet fully busy, it cannot stop the price of what that compute produces from falling. I believe these two “branches” frame the uncertainty pretty well - in the demand-heavy branch, commitments resell while margin pressure builds; in the demand-soft branch, the heterogeneous frontier essentially crushes pricing. These are pressure mechanisms, not guaranteed outcomes - but a provider locked to one hardware stack is poorly hedged in either branch - I don’t see how that position wins both - and hedges are precisely what a portable spot market and a futures curve exist to provide. As I understand his case, it holds while that scarcity persists - and spot rates above contract are exactly what it looks like today. However, my position is that when the heterogeneous capacity starts landing, that logic flips - the same contract roll-off he is counting on becomes the mechanism that reprices compute downward. And to the objection that today’s heterogeneous fleets are non-existent - I would state that prices are set at the margin, not by the average - but only once enough workloads can actually access, qualify and substitute onto that capacity at meaningful volume. From there, qualified heterogeneous capacity at meaningful scale can begin resetting marginal prices - portability is one of the prerequisites that qualifies it, alongside fidelity, volume, service levels and merchant availability - and private capacity can increasingly become merchant capacity at that price. Google now sells TPUs outward, Trainium is rapidly scaling, Cerebras and Trainium are being paired, NVIDIA licensed Groq’s architecture inward, and FIG 03 maps exactly this arrival schedule.

AI ops is the real moat

The third argument is the least discussed, but one of the strongest in my view as the scaled evidence is already in plain sight. That argument is that serving excellence is the actual moat - open-source engines are free, the last mile matters most - and a heterogeneous world makes that last mile harder and therefore more valuable, the way Linux being free made EC2 a money machine. If you look at what EC2 actually is - Amazon runs Intel, AMD and its own Graviton behind one control, procurement and operational layer - distinct instance classes, not perfect interchangeability - and then they basically arbitrages them incredibly well. EC2 is the closest realized version of the portability thesis at the CPU layer, and its margin comes from owning the fungible compute substrate and the operational excellence on top - not from captivity to one hardware vendor. I’m arguing this is exactly the position that inference platforms should want - become the AWS of accelerators. What portability displaces is not the serving margin - it is the captivity rent - and a platform whose stack cannot route to the silicon that enables the fastest tokens is defending the wrong thing. Portability removes captivity rent; it does not remove the margin earned by running a superior platform.

Portability needs to be vendor neutral

The fourth argument is that any portability layer that is successful ultimately becomes the new lock-in - our new world has just moved the concentration, not removed it, and I agree with this point. A portability layer only catalyzes if it has open interfaces, reproducible tests a third party can run, flexible licensing rights that can survive any single vendor, and enables a global ecosystem to implement and replace their existing layer. That is the standard I have always believed in, and it’s critical that this is the standard we drive in the future.

Where this leaves the market

All in all, lets combine these conclucsions and create a simple example that forecasts the world they create.

Let’s imagine that we take the “captive financing” from the Every GPU loan is really a software loan essay - those mechanics look like a sub-five-year repayment schedule, secured debt near 8.8% (the 9.75% unsecured debt sits above it), and compute that is roughly half idle - and we replace each term with a more liquid-market case using the learnings from aviation: lower financing costs, a longer earning life supported by residual demand and remarketing, and utilization lifted by workload routing. Under those terms, the liquid-market required recovery falls from $1.60 to roughly $0.77 per GPU-hour - a scenario, not an outcome portability alone demonstrates.[14] Cross-silicon competition then presses the price further down because a clearing level would need a supply-and-demand model. To make that clear and translate that into token economics: the production serving rate documented in note 5 - roughly 300 to 400 tokens per second on a trillion-parameter-class model - is a replica rate (I couldn’t find the full config details) so I won’t convert it into absolute dollars per million tokens. Holding workload and throughput constant though, the ratio is independent of replica size: $1.60 against $0.77 of required recovery means the liquid-market case roughly halves the capital-recovery component of every million tokens served.

To be clear - I am not estimating the percentage reduction in total break-even token cost, since power, networking and operations shares vary widely by deployment. Portability is not the sole source of the earning tail - CoreWeave is generating one inside NVIDIA’s ecosystem today. Its role is to broaden that tail - expanding the qualified hardware, workloads, operators and exit buyers available to the asset, and making its value less dependent on one vendor’s software stack. And the chart’s real point is to emphasize that the spread compression is worth six cents, and a longer amortization window is worth forty-nine. The interest rate matters less than the amortization window that portability could help lenders support. If we strip the extension out - keep the 4.5-year window, with the better financing at 75% utilization - and the requirement only reaches about $1.13 - so the window is the dominant term in this capital-recovery model. A longer window also has to be real: the economic condition where the older machine’s lower acquisition cost and routing value outweighs a newer chip’s power efficiency and performance - the same old checkpoint can be cheaper to run on new silicon, so residual demand has to be factored in. Portability helps lenders underwrite a longer life; it does not by itself create one. Essentially, the simple flow is that financing terms set the required hourly rate, the hourly rate sets the operator’s break-even token cost, and portability can improve financing, utilization and hardware choice at once - while adding the validation and coordination costs above.

FIG. 05   Panel A: required capital recovery per GPU-hour repriced one factor at a time to an illustrative liquid-market case, with a one-variable useful-life sensitivity (5.5 / 7 / 8 years at fixed financing and utilization). Panel B: cross-silicon competition as direction, not a point estimate. Illustrative model on a secured captive comparator; assumptions in note 14.

Token demand is compounding faster than any industrial input I have ever seen, token prices are collapsing underneath it (and are likely to continue falling), and financing currently protects itself from that deflation through short amortization, customer contracts and credit wrappers, rather than underwriting against a transparent forward token-cost curve. My first essay argued that making silicon programmable and making it financeable are the same problem, and this essay argues that once a portable layer gives the market a visible cost curve, the long-contracted GPU market gets repriced significantly. CoreWeave shows that software continuity and operational scale can already create reuse, recontracting and liquidity inside one hardware ecosystem - portability’s larger promise is to extend that liquidity across vendors and expand the pool of workloads, operators and buyers available to each machine. Once enough heterogeneous capacity becomes production-ready, merchant and reachable through portable software, it becomes qualified substitute supply and begins resetting the marginal cost of inference. When portability then makes that capacity easy to measure, easy to reach, and possible to hedge, falling token prices stop being a threat and become an input variable into a loan. At that point, what gets financed could shift from the chip itself to the work it produces - a verified unit of inference, the token.

Footnotes

1. The Bank for International Settlements gap measurement per I. Aldasoro, S. Doerr and D. Rees, “Financing the AI boom,” BIS Bulletin No. 120, January 2026 - https://www.bis.org/publ/bisbull120.pdf . Tokens served per Sundar Pichai’s Google I/O keynotes: 9.7 trillion per month (May 2024), roughly 480 trillion (May 2025), over 3.2 quadrillion (May 2026) - https://blog.google/innovation-and-ai/sundar-pichai-io-2026/ . Alphabet capital expenditures: $52.5 billion (2024) and $91.4 billion (2025) per the FY2025 10-K - https://www.sec.gov/Archives/edgar/data/1652044/000165204426000018/goog-20251231.htm ; 2026 guidance of $175-185 billion in February (per CNBC - https://www.cnbc.com/2026/02/04/alphabet-resets-the-bar-for-ai-infrastructure-spending.html ), raised to $180-190 billion at Q1, then to 195-205billiononJuly22,2026;thechartusesthe~200 billion midpoint. Capex per million tokens = calendar-year capex ÷ (May monthly tokens × 12): $52.5B ÷ 116T ≈ $451; $91.4B ÷ 5,760T ≈ 16;~200B ÷ 38,400T ≈ $5. This is an intensity ratio, not a unit cost: capex leads usage, includes training and non-serving infrastructure, and Google’s token count spans all surfaces. Constant-capability market price per note 4, anchored at GPT-4-class output of roughly $60 per million (March 2023) declining ~10x per year. Methodology for the capex-intensity series: calendar-year Alphabet capex divided by twelve times the May keynote monthly token run-rate; Google’s token definitions and scope may not be consistent across years, so treat it as directional intensity, not unit cost - halving or doubling the run-rate assumption moves the 2026 figure between roughly $2.50 and $10 per million tokens.

2. CoreWeave Q1 2026 earnings release (SEC-filed), including total debt of $24.86 billion and interest expense of $536 million on revenue of $2,078 million - https://www.sec.gov/Archives/edgar/data/0001769628/000176962826000220/coreweave1q26earningspress.htm . The DDTL 3.0 covenant amendment (effective December 31, 2025) per CoreWeave Form 8-K filed January 2, 2026, following customer delivery-timing changes - the amendment reduced the minimum liquidity requirement, postponed initial covenant-testing dates and expanded equity-cure rights - https://www.sec.gov/Archives/edgar/data/1769628/000176962826000003/crwv-20251231.htm . CoreWeave Q2 2026 results: revenue of $2.575 billion with net interest expense of $640 million (roughly 25 cents per revenue dollar) and depreciation and amortization of $1.393 billion; adjusted EBITDA margin of 59% against a 5% adjusted operating margin; a roughly $104 billion revenue backlog with more than $25 billion of additional commitments after quarter-end; a roughly 25% price increase across SKUs in July; expected contribution margins on newly signed contracts five to ten percentage points above prior quarters; prior-generation ASPs at or above year-earlier levels with Ampere and Hopper fleets described as largely sold out; an A100 capacity contract extending into 2029 at what management called an attractive price; managed-inference booked ARR growing from roughly $1 million to more than $100 million within months, targeting at least $250 million by year-end and framed as redeploying GPUs coming off contract; and over $400 million of ARR from storage, CPU, networking and software - https://investors.coreweave.com/news/news-details/2026/CoreWeave-Reports-Strong-Second-Quarter-2026-Results/default.aspx . DDTL 5.5: a $2.6 billion facility of roughly five-year money against customer contracts averaging about three years, with renewal or re-leasing of the capacity contemplated - reflecting lender willingness to underwrite renewal risk - https://investors.coreweave.com/news/news-details/2026/CoreWeave-Closes-2-6-Billion-Loan-Facility-Expanding-Financing-Flexibility-for-AI-Infrastructure/default.aspx

3. Representative providers in this layer include Fireworks AI, Baseten and Together AI. Procurement and margin figures per Sacra research on Fireworks AI (reserved capacity contracted on longer commitments; ~50% blended gross margins with a stated 60% target via utilization improvements; revenue “likely concentrated among a smaller number of large production deployments” against a customer base of 10,000+) - https://sacra.com/research/fireworks-ai/ ; Baseten $300 million Series E at $5 billion, February 2026, per the same. The margin gradient described in the text - expansive on smaller list-priced customers, compressing sharply on individually negotiated large deployments - is stated as market structure, not as a disclosed figure.

4. Guido Appenzeller, “Welcome to LLMflation,” Andreessen Horowitz, November 2024 (10x per year for equivalent performance; GPT-3-class from $60 to $0.06 per million tokens, 2021-2024) - https://a16z.com/llmflation-llm-inference-cost/ ; Cottier, Snodin, Owen and Adamczewski, “LLM inference prices have fallen rapidly but unequally across tasks,” Epoch AI (9x-900x per year by task) - https://epoch.ai/data-insights/llm-inference-price-trends . Current-frontier pricing: MiniMax M3 at 0.30/1.20 per million tokens in/out, June 2026 - https://developer.puter.com/tutorials/minimax-api-pricing/ ; GLM 5.2 at 1.40/4.40 and aggregate scores alongside or above closed flagships per BenchLM’s July 2026 index (M3 and Claude Opus 4.5 both scoring 71; GLM 5.2 at 81 vs Gemini 3 Pro at 79) - https://benchlm.ai/llm-pricing ; Moonshot Kimi K3 at 3/15, released July 16, 2026, per Artificial Analysis - https://artificialanalysis.ai/models/kimi-k3 . Per-task cost and verbosity per Artificial Analysis’s Intelligence Index: Kimi K3 at $0.94 per task on ~130M output tokens (double the 63M median; 21% fewer than K2.6), GPT-5.6 Sol at $1.04, Claude Opus 4.8 at 1.80,GLM-5.2at~0.32-0.47 - https://x.com/ArtificialAnlys/status/2077832874183860404 . Aggregate benchmark indices are directional, not definitive. The ~23% annual GPU-hour decay is the author’s synthesis of US-region one-year reserved H100 (SXM) rates across 2023-2026 from the SemiAnalysis rental index and Silicon Data tracking (the first essay’s FIG 01); spot rates and other regions differ. The constant-capability envelope in FIG 01 and FIG 02 is a single illustrative construction - a fixed reference task at a defined quality threshold, anchored at GPT-4-class output in March 2023 - not a measured index; the a16z GPT-3-class arc and the current-model per-task examples are observations within the same envelope at different dates.

5. Latent Space: The AI Engineer Podcast, “The Inference Engineering Masterclass,” with Baseten’s Philip Kiely and Ali Taha (hosts swyx and Vibhu Sapra), August 2026 - https://www.latent.space/p/inference-eng . Kiely describes a production loop in which GLM 5.2 profiles Baseten’s serving stack and writes replacement SGLang GPU kernels that then serve GLM 5.2 itself, validated by a second trace with a human pulling the image; the same conversation documents the roughly 10x gap between off-the-shelf and production serving (30-50 versus 300-400 tokens per second on a trillion-parameter model), quantization fidelity verified at the logit level via KL divergence, identical weights behaving differently across clusters, and KV-cache movement and interconnect as the next bottleneck.

6. Heterogeneous supply arriving 2026-2030: Google Q1 2026 earnings call - TPU hardware sales with the bulk of revenue expected in 2027, per the company’s Q1 2026 earnings call; Blackstone-Google TPU cloud, first 500 MW targeted for 2027; Anthropic agreement with Google and Broadcom for 3.5 GW of capacity starting 2027, per Futuriom - https://www.futuriom.com/articles/news/google-and-blackstone-to-create-tpu-cloud-service/2026/05 ; AMD-OpenAI agreement, October 2025: 6 GW of MI-series capacity across generations, first gigawatt of MI450 in the second half of 2026; Cerebras-OpenAI Master Relationship Agreement, December 2025: 750 MW of wafer-scale inference capacity in tranches 2026-2028, option for a further 1.25 GW, and $24.6 billion of remaining performance obligations per the S-1, via SemiAnalysis - https://newsletter.semianalysis.com/p/cerebras-faster-tokens-please ; Qualcomm AI200 and AI250 rack-scale inference systems, commercially available 2026 and 2027 on an annual product cadence, with Humain deploying 200 MW from 2026, per Data Center Dynamics - https://www.datacenterdynamics.com/en/news/qualcomm-launches-ai200-and-ai250-chip-offering-targeting-inferencing-workloads-at-rack-scale/ ; NVIDIA-Groq: roughly $20 billion non-exclusive license of Groq’s LPU inference technology plus the hiring of founder Jonathan Ross and senior team, December 2025, with Groq pivoting to an inference cloud on a $650 million raise, per TechCrunch - https://techcrunch.com/2026/06/22/ai-chipmaker-groq-confirms-650m-raise-re-staffs-after-nvidias-20b-not-acqui-hire-deal/ and Bloomberg - https://www.bloomberg.com/news/articles/2026-06-22/groq-raises-650-million-to-help-startup-pivot-after-nvidia-deal ; Meta TPU agreement reported 2026, per The Next Web - https://thenextweb.com/news/google-blackstone-tpu-cloud-joint-venture-5bn ; NVIDIA Rubin on the publicly committed one-year cadence. Deployment windows in FIG 03 reflect these announced schedules by confidence category. Not forecasts. AWS Trainium: custom-silicon annualized run rate above $20 billion with $225 billion in multi-year commitments per Amazon’s Q1 2026 earnings call (April 29, 2026); OpenAI committed roughly 2 GW of Trainium (ramping 2027) and Anthropic up to 5 GW across Trainium generations, per Data Center Dynamics - https://www.datacenterdynamics.com/en/news/aws-partners-with-big-chip-co-cerebras-for-ai-inference-disaggregation/ . AWS-Cerebras disaggregated pairing (Trainium prefill, CS-3 decode, EFA interconnect, on Bedrock), announced March 13, 2026 - https://www.cerebras.ai/press-release/awscollaboration

7. “Adequate” is measurable, not rhetorical: fidelity to the reference model attested at the logit level - KL divergence between served and reference output distributions within tolerance - at the workload’s latency class and context length. The attestation infrastructure exists in embryo: model developers publish vendor-verification suites (Moonshot’s K2 Vendor Verifier is the template), and serving providers publish logit-fidelity audits of their own quantizations (note 5). It is the same infrastructure a token forward curve needs for grading.

8. Liang Wenfeng investor-call transcript, recorded May 20 and compiled July 16, 2026, pp. 9-11, 13 and 31-35. Liang describes AI development as a cumulative staircase from language models to chain of thought, agents and continuous learning. He treats self-iteration as a consequence of continuous learning and embodiment as downstream of self-iteration, rather than as separate intermediate rungs. The circulated document was generated through automated speech recognition, cleaned with AI, and uses speaker labels inferred from context.

9. Disaggregated serving: Zhong et al, “DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving,” OSDI 2024, arXiv:2401.09670 (up to 7.4x more requests within latency targets) - https://arxiv.org/abs/2401.09670 ; Patel et al, “Splitwise: Efficient Generative LLM Inference Using Phase Splitting,” ISCA 2024, arXiv:2311.18677 (2.35x throughput at the same cost and power) - https://arxiv.org/abs/2311.18677 ; Qin et al, “Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving,” FAST 2025 best paper, arXiv:2407.00079 (Moonshot’s production Kimi platform; 75% more requests handled, up to 525% throughput in long-context scenarios) - https://arxiv.org/abs/2407.00079 ; NVIDIA Dynamo disaggregated inference framework (GTC, March 2025); Rubin CPX prefill-specialized GPU (September 2025; 30 PFLOPS NVFP4, 128 GB GDDR7, no NVLink; Vera Rubin NVL144 CPX with 144 CPX + 144 Rubin GPUs shipping end-2026; NVIDIA projects $5 billion of token revenue per $100 million invested) per SemiAnalysis - https://newsletter.semianalysis.com/p/another-giant-leap-the-rubin-cpx-specialized-accelerator-rack and Futurum - https://futurumgroup.com/insights/nvidias-new-rubin-cpx-targets-future-of-large-scale-inference/ ; NVIDIA’s post-Groq inference architecture (GPUs for prefill, SRAM LPUs for decode) per Spheron - https://www.spheron.network/blog/nvidia-rubin-cpx-long-context-inference/ Practitioner consensus increasingly locates the next order-of-magnitude unlock in interconnect and state movement rather than raw FLOPs - see the serving-stack accounts in note 5. The cross-vendor version is now shipping: AWS pairs Trainium prefill with Cerebras CS-3 decode over EFA as a Bedrock service (note 6).

10. Convergence from both directions: Together AI’s co-built GPU cluster and data center program with Hypertec/5C (36,000 NVIDIA GB200 NVL72 GPUs announced November 2024; secured capacity for 100,000+ GPUs; up-to-2 GW European AI-factory alliance announced June 2025) - https://www.prnewswire.com/news-releases/together-ai-and-hypertec-cloud-join-forces-to-co-build-turbocharged-nvidia-gb200-cluster-of-36k-blackwell-gpus-302308267.html ; https://hypertec.com/hypertec-expands-into-europe-with-5c-and-together-ai-a-strategic-ai-infrastructure-alliance-driving-up-to-5-billion-in-private-investments/ ; CoreWeave’s acquisition of Weights & Biases (announced March 2025), an operator buying up the stack into developer tooling and inference.

11. Aircraft leasing’s unbundling arc: 2.4% of the global fleet leased in 1980; the leased share of the world’s passenger jet fleet crossed 50% for the first time in 2021 per Cirium; 51.3% (12,000+ aircraft) by late 2024 per Air Lease founder Steven Udvar-Hazy - https://aviationweek.com/air-transport/airlines-lessors/proportion-leased-aircraft-stabilize-around-50-lessor-says ; CAPA, “Aircraft leasing in equilibrium at just over half the world fleet,” February 2024 - https://centreforaviation.com/analysis/reports/aircraft-leasing-in-equilibrium-at-just-over-half-the-world-fleet-675212

12. “Open Weights and American AI Leadership,” open letter, July 24, 2026; 25 signatories including NVIDIA, Meta, Microsoft, Andreessen Horowitz, IBM, Mistral, Hugging Face, Mozilla, the Linux Foundation, Palantir, Perplexity and Y Combinator - https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/ ; shared by Jensen Huang as his first post on X (“The world needs both frontier closed models and frontier open models”), coverage per Decrypt - https://decrypt.co/374282/nvidia-meta-microsoft-washington-dont-kill-open-source-ai

13. Gavin Baker (Atreides Management), X thread, July 28, 2026 - https://x.com/GavinSBaker/status/2082166566280642676 - arguing spot GPU rental rates at least 2x contracted rates imply hyperscaler underearning; hyperscale operating-cash-flow growth accelerating from roughly 31% (Q1 2026) toward 50% (Q2 2026) on his estimates; CY2028 capex of $1.5 to $2.2 trillion against $1.3 to $1.4 trillion of consensus operating cash flow; NVIDIA and Broadcom guarantees framed as credit “wrappers.” Cited exhibits include Wells Fargo hyperscaler CDS and leverage charts and the BofA hyperscaler-versus-semiconductor free-cash-flow series.

14. Author’s illustrative repricing model. Required capital recovery per GPU-hour: $30,000 H100 system-allocated, straight-lined over the amortization term, plus interest on a $12,000 average financed balance (80% advance), divided by utilized hours. The captive comparator deliberately uses a secured DDTL 5.0-class rate of roughly 8.8% - not CoreWeave’s 9.75% senior unsecured print, which is reported in the first essay for context but is not a like-for-like collateral rate. Captive: 4.5-year amortization (per Moody’s on DDTL 4.0), 55% utilization, 8.8% - 1.60perGPU-hourofrequiredcapitalrecovery.Isolatedrecoverysteps:financingcompressed250bpto6.3%(-0.06); amortization extended to seven years against the earning tail (-0.49);utilizationto75%viarouting(-0.28) - an illustrative liquid-market required recovery of $0.77. The $30,000 system-allocated basis and 55% captive utilization are chosen illustrative inputs, not sourced observations. One-variable useful-life sensitivity, holding financing (6.3%) and utilization (75%) constant: 5.5 years roughly $0.95, 7 years $0.77, 8 years roughly $0.69. Composite bear/bull cases that also vary financing and utilization together span roughly $1.11 to $0.61 and are labeled as composites, not life sensitivity. The economic condition behind any life extension is a crossover - the older machine’s lower acquisition cost and routing value must outweigh newer silicon’s power efficiency and performance for residual demand to exist. Cross-silicon competition is shown directionally only - it pressures the blended fleet and replacement cost that clears around the floor, but a market-clearing level would require a supply-and-demand model this note does not attempt. A no-life-extension case - the 4.5-year window with 250bp compression and 75% utilization - reaches only about $1.13, isolating the amortization window as the dominant term. The note 5 serving rate (300-400 tokens per second, trillion-parameter class) is a replica rate whose GPU count, batching, precision and concurrency the citation does not specify, so this note reports the ratio only: holding workload and throughput constant, replica configuration divides captive and portable recovery equally, leaving the roughly-half reduction in the capital-recovery component intact; no absolute dollars-per-million-tokens figure is claimed. The seven-year amortization is supported qualitatively by workload ossification: enterprises pin validated checkpoints for years, keeping older silicon in production service. CoreWeave’s 2026 A100 recontracting and DDTL 5.5 (note 2) provide qualitative evidence for residual demand and lender willingness to underwrite beyond an initial customer contract; neither validates the seven-year assumption or the $0.77 point estimate. Operator-level serving-efficiency variance - note 5 documents a roughly 10x gap between off-the-shelf and production serving on the same model class - is a separate software-execution effect and is deliberately not counted anywhere in this model. Gross of power, networking and operations; not a forecast.

Notes above give full source attributions. Figures indicative; floating tranches shown at approximate all-in rates.

Read More
Interviews Tim Davis Interviews Tim Davis

EE Times: Why Qualcomm Bought An Open AI Software Stack

In an exclusive interview with EE Times, Modular co-founder Tim Davis breaks down why Qualcomm acquired the startup for nearly $4 billion and how they plan to build an open, heterogeneous future for AI compute.

I spoke with Sally Ward-Foxton at EE Times about Qualcomm’s acquisition of Modular and what it means.

You can find the full article here.

In an exclusive interview with EE Times, Modular co-founder Tim Davis breaks down why Qualcomm acquired the startup for nearly $4 billion and how they plan to build an open, heterogeneous future for AI compute.

Sally Ward-Foxton, EE Times

When Qualcomm acquired AI software infrastructure startup Modular for close to $4 billion, it appeared to be a strange fit. Modular has built a hardware-agnostic software stack designed to abstract away the underlying silicon from developers, enabling portability. Qualcomm sells proprietary silicon, so portability wouldn’t necessarily be in its interests. But according to Modular cofounder and president Tim Davis, the acquisition represents a shared belief that the future of AI compute is heterogeneous, and that software is the key to deployment of non-GPU hardware for AI inference.

One of Modular’s original theses, Davis said, is that the future of AI compute would certainly be heterogeneous. Their other theses—that inference was going to be huge, that superintelligence would need both data center and edge hardware, and that software was holding everything back—have all come true.

“If you tie those four theses together across the ecosystem, and then if we were to consider a partner to help us achieve our mission, who ticks all the boxes and who will enable us to have global distribution while still maintaining an open, neutral platform for the world?” Davis told EE Times in an exclusive interview.

Qualcomm has an enormous hardware footprint, from edge to cloud, so it’s certainly aware of the challenges of running AI on heterogeneous silicon, Davis said. “We felt confident in the Qualcomm team and Cristiano [Amon, Qualcomm CEO] and the other leaders at Qualcomm,” Davis said. “The ability to get extreme distribution much faster by having a backer of that scale was an exciting proposition.”

Velocity is critical if Modular wants to become a new open standard for the world, both in the data center and at the edge.

“The folks at Qualcomm are very excited about maintaining an open, neutral standard, being able to continue to work across all silicon,” Davis said. “The heterogeneous nature of compute is such that for everything to be successful, we must achieve that mission.”

Importantly, Modular will remain an independent entity—the team will be maintained and the mission preserved, Davis said.

“We have a shared understanding that we want an open developer ecosystem,” he said. “To have an open developer ecosystem that drives AI globally, you need an open set of infrastructure that developers are excited about, otherwise there won’t be an ecosystem.”

Mojo and Max

Modular’s key products are Mojo, a kernel programming language that is silicon-agnostic, and Max, an open modelling and serving framework that enables heterogeneous scaling. Both Mojo and Max are open source.

Mojo is designed to make it easy to write core infrastructure that’s natively multi-silicon, covering CPU-ASIC, heterogeneous GPU, and cloud-edge silicon use cases. Mojo 1.0, a stable, production-ready version of the language that has been under development for 4.5 years, will launch shortly.

“Everyone wants a multi-silicon future, for a range of reasons,” Davis said. “A lot of the agentic workflow is now CPU-bound, and we’re getting disaggregated inference becoming a big thing.”

Cloud management and orchestration is still an open problem, Davis said, which requires a software solution because of the shift from selling GPU-hours to selling tokens.

“We want a multi-silicon world where the user doesn’t have to care where the tokens are being generated,” he said. “We made a radical statement four years ago: Developers do not care about the hardware. That is more true today than ever. The industry is focused on tokens per second.”

Max builds on Mojo and is similarly multi-silicon by design. Max is already suitable for multi-modal inputs and outputs, where today’s serving infrastructure is predominantly text-optimized. Projects like vLLM and SGLang have forks that are image and video-optimized, but this is already too much fragmentation, in Davis’ opinion.

Modular has done a lot of work on management, observability, MLOps, and cloud optimizations for Max, Davis said. Max is due to go into general availability in a few weeks.

“We will, over time, add training,” he said. “Then you’ll have a stack that is multi-modal, multi-silicon, and enables you to go from training to inference, from edge to cloud. That’s the dream.”

Alternative hardware

Software infrastructure has been holding back deployments of non-GPU hardware at scale so far, he said. Neoclouds finance most of their GPU cluster buildouts using debt instruments, which is fine for Nvidia or AMD hardware, (they have a known depreciation curve and the assets are liquid, which means they can be sold to someone else), but for alternative silicon there are few financing options today, Davis said. Neoclouds are having to fund clusters with equity, which is extremely expensive and limits fungibility across compute.

“If you have a multi-silicon software layer that makes the fungibility of an asset easier, everyone around the world starts to say that the existing compute economy has overpriced software on a certain provider, and it doesn’t mean the underlying compute asset has changed, it just means that fungibility across compute helps rewrite the financial equation for why more silicon can get into clouds, which really matters,” Davis said.

Qualcomm has realized how big an issue software is, knows that a consistent platform is needed to reduce the burden on developers, and understands that an open ecosystem is key to making that happen, Davis said.

Modular’s customers so far are the most sophisticated enterprises and genAI startups, plus Neoclouds and token factories.

“We have customers at every level of the stack,” Davis said. “We have hardware customers—people who want silicon brought up—we have neoclouds who are buying that hardware and want to install it in their token factory so they can make bigger margins, frontier model labs who want to sell to enterprises via API, and we have enterprises who don’t care where the tokens come from.”  

Frontier labs increasingly like multi-silicon architectures as they become total cost of ownership (TCO)-driven, Davis said.

“Frontier labs want us to enable them to develop on our open infrastructure and give them multi-silicon diversity, which is great for their business, because now they can work with everyone to scale it across multiple layers of the stack, which is really important,” he said. “They then become part of the open ecosystem we’re creating, helping improve our infrastructure but actually improving it for everyone.”

AI kernel generation

EE Times has spoken to a number of chip companies recently that have augmented or even replaced parts of their compilers with AI assistants or agents. Writing a performant compiler for novel hardware has been a huge barrier to entry for these companies and they are embracing AI code generation in their stacks, perhaps making a Claude-based agent, teaching it about the supplier’s hardware, and then letting it write or convert code at the lowest levels.

Davis is more cautious. Compilers need to be deterministic, he said, noting that AI almost by definition doesn’t produce deterministic outputs, and that mechanistic interpretability is still a huge area of research.

“AI code generation is very exciting, but I view it as an inverse pyramid—it’s amazing at building prototypes, it’s amazing at certain application layer stuff, but the lower you go, the more determinism you want,” he said.

Every enterprise will have to decide how much non-determinism they are willing to accept in their code base. For some organizations Modular works with, including financial institutions, any is too much.

“The question is: are you willing to take a risk on production workloads with production customers on vibe-coded stuff at the end of the day?” Davis said.

Simply writing kernels that work is also not enough, Davis said—they have to be comparable or better than CUDA.

“Compute portability is still about TCO,” he said. “The fact that AI can write kernels is great, and it’s exciting, but you need to have near state-of-the-art performance for the TCO to work on most alternative silicon today.”

For many reasons, AI code generation still needs a human in the loop, Davis said.

“I still think there are people in the world that know that part of the stack incredibly well,” he said. “Above that, we need people who are experts in techniques like speculative decoding… there’s a lot of sophistication in that part of the stack too. And then above that, you need cloud infrastructure to actually orchestrate it all. So it is a whole-stack problem.” 

AI-generated kernels aren’t ready for production deployment today, evidenced by existing limitations on commercial stacks, Davis said. The future will need a core substrate to build on, namely, Mojo and Max, he added.

“Once you have a unified programming model at the base that can expose the core intrinsics of your silicon, you get all the AI framework and the serving and everything else just out of the box that you can go and use. That is not true [elsewhere] today, you can’t just go and vibe-code that. It’s a hard systems programming problem to do that and then to tie it all together.”

Modular’s team is excited about convincing the world that the future is heterogeneous.

“It’s inevitable, if you look at history, that we will get there,” Davis said. “But how we get there is an open question. Our strong belief is if we can help all of these companies get there faster by not having to worry about the software as much, that’s going to be wonderful for the world.”

Read More
Essays Tim Davis Essays Tim Davis

Every GPU loan is really a software loan

AI’s debt boom is secured by GPUs but priced off contracts. The underlying assets are the GPUs, but the financing depends on the contracts, guarantees and operating rights wrapped around them.

A multi-trillion-dollar AI debt buildout is underway, and I wanted to understand how the market is actually funding and pricing it. What I found is that a GPU loan is only as good as the next operator’s ability to run the software stack.

One of the largest debt buildouts of the decade is in AI hardware, and I wanted to dig into the financing of it to understand how the market is pricing it. One of the most basic initial questions I had was: if a lender of GPU debt has to repossess the GPUs - what exactly do they end up holding? I discovered that the answer has far less to do with the silicon than with the software, because software determines whether the next owner can fundamentally put those machines back to work. I was interested in the debt markets view, because the flow of funds is always an interesting story as that’s where the underlying risk is forced into a price.

This is not meant to be an essay about whether AI debt is a bubble, rather, it’s an essay about what makes compute financeable at all, and my take from all the research below is that it’s mostly software, not silicon.

Across this essay, I use “portability” to mean three related things:

  1. Workload portability means a model can move across different silicon with bounded cost in performance, correctness, and engineering effort.

  2. Operational transferability means a qualified third party can take over a running cluster and keep it earning.

  3. Asset liquidity means the hardware itself can find a new operator or buyer at an observable price.

The argument of this essay is that a causal chain exists that the market isn’t appreciating or understanding, and here’s a simple summary of how the above three points relate - the first makes the second cheap, the second is what makes the third real, and the third is what the debt markets actually price.

I'll also refer to “fungibility” a lot, and here I do not mean every accelerator is identical or that every workload belongs on every chip. I simply mean the interchangeability inside a machine’s “capability window” - that is, whether that workload can move among the machines capable of serving it with predictable performance, cost and correctness within their usable life. My view is that compute will become fungible first where the switching cost is lowest - inference, batch processing, much of fine-tuning - long before it reaches frontier training, where a failure can cost weeks and where the collective communication patterns bind you even more to a single hardware vendor.

The numbers underneath the buildout

The magnitude of what is being borrowed against AI hardware to finance the future of compute has reached unprecedented levels, and I don’t think most people appreciate the sheer scale of it. AI-related companies and projects raised at least $200 billion of new debt in 2025 alone - a figure that undercounts the private deals - and that lands on top of a much larger instrument - by October 2025, the outstanding bonds of AI-tied issuers totaled $1.2 trillion, the single largest sector of the US investment-grade market at roughly 14%, larger than American banks - though not all of it was raised specifically to finance AI infrastructure.[1] The five largest hyperscalers issued $121 billion in new bonds in 2025 alone, more than four times their average annual issuance over the prior five years, and the four largest have guided to roughly $725 billion of combined capital expenditure in 2026 - up nearly 80% on what they spent the year before, with Alphabet alone raising its range to $195 to $205 billion on July 22, 2026.[2]

In my research, I learnt that underneath all this financing sits a newer structure - securitized data center debt, which JPMorgan’s CMBS research team projects will reach $30 to $40 billion a year in 2026 and 2027, up from roughly $27 billion in 2025.[3] And the pace is accelerating - AI-related bond issuance topped $250 billion by the middle of 2026, even as the first signs of strain appeared - Amazon’s latest $25 billion sale seemingly met a cool reception, S&P cut Oracle to the last rung of investment grade over its AI spending, and Alphabet and Meta both paid real premiums over their prior deals just to access the market.[4] The Bank for International Settlements put a frame around all of this in January - the boom has structurally shifted from cash-flow financing to debt - free cash flow at the major AI firms now lags their capital expenditure outright - with private credit to AI companies growing from near zero to over $200 billion, roughly 8% of all outstanding private credit, and projected by the BIS to reach $300 to $600 billion by 2030.[5]

So where the hell is all this money going? At one end of this sit the neoclouds, whose entire business model is fundamentally borrowing against GPUs. Critically, these businesses don’t have a separate foundational core business like the major hyperscalers - for example, Google with Ads, Android or GCP, or Amazon with Shopping or AWS, or Microsoft with Office - that provides a strong and healthy balance sheet to feed the capital expenditure. CoreWeave, really the “template” for the category, has grown its total debt to roughly $25 billion - up from under $8 billion at the end of 2024 - across a series of delayed draw term loans collateralized by NVIDIA hardware and the customer contracts attached to it. I dug into their financials and found that when CoreWeave raised its first facility in August 2023, it was the first time H100s had ever been used as loan security.[6] Interestingly, every one of these structures awkwardly embeds an assumption about hardware residual values - that is, they need to determine what the hardware is worth in year three, year five, year seven, to a lender who may one day have to repossess it and find someone else to own it, and run it. That residual-value assumption is where a really interesting narrative lives, and I wanted to double down on that as I believe it deserves much closer scrutiny.

Someone else is setting the clock

A few close friends and I discuss (and argue! haha) regularly about compute and the future of AI. One of the bear case discussions we have around compute is as follows - software portability threatens compute’s underlying value - that is, if workloads can move freely across silicon, the premium NVIDIA hardware commands will erode and the loans written against it will ultimately sour, or at least be devalued. It’s a great hypothetical and I wanted to illustrate my position around why I think it gets the causality backwards.

I took a look at what actually happened to Hopper over the last few years. H100 rental pricing peaked around $8 per GPU-hour during the scarcity of 2023, with some on-demand rates exceeding $12. By mid-2024 the same capacity rented for $2 to $3. One-year reserved pricing bottomed at roughly $1.70 in October 2025 before recovering to about $2.35 by March 2026 as inference demand surged, settling around $2.25 by mid-2026.[7] Interestingly, the price of a three-year-old H100 system today spans from roughly 45% of new price (hashrateindex’s used-market tracking) to roughly 69% (Silicon Data’s index).[8] I wanted to compare that to a well-known, well-understood industry like aircraft, a staple of debt financing (I went into this knowing absolutely nothing about aircraft financing mind you!) Aircraft appraisal is imperfect, but it rests on a strong set of standardized principles - maintenance records, known configurations, decades of transaction history, and a far deeper remarketing market than compute has today. Against that baseline, GPUs have an incredibly wide pricing window - spanning 45 to 69% at the same age - which is fundamentally what an asset without a true liquidity layer looks like when you try to value it. I attempted to graph this below:

FIG. 01   H100 GPU-hour pricing, 2023-2026. Sources: SemiAnalysis, Silicon Data, Spheron.

The obvious question to ask after looking at the above graph is: what caused Hopper values to fall and slowly rise? Hopper repriced sharply as Blackwell began shipping - so the simplest conclusion is NVIDIA’s own release cadence, which they have publicly committed to now as a one-year rhythm. Put simply, when your hardware is locked to a single vendor, you don't get to decide how fast it loses value - they do. More directly, the GPU’s collateral value is a function of NVIDIA’s roadmap and is enormously concentrated by it.

In my research, it seems the industry can't even agree on the single most important input to the model. In the same quarter, Amazon shortened the assumed useful life of a subset of its servers and network equipment from six years to five, explicitly citing the accelerating pace of AI hardware - and took $920 million of accelerated depreciation retiring equipment early - while Meta extended the lives of certain servers and network assets to 5.5 years, cutting its 2025 depreciation expense by roughly $2.9 billion.[9] Neither company breaks these estimates out by chip type - they are fleet-level judgments, not GPU specifics (which sucks) - but the direction is the point - same reporting period, the same asset category, and opposite conclusions that are worth billions. The point is not that one company is wrong; it's that no stable consensus exists on useful life even at the fleet level, and there is little transparency into it either. When the most sophisticated infrastructure operators on the planet cannot agree there, GPU-specific residual value is even less settled - which is precisely the condition under which lenders demand something else to underwrite against.

Unsurprisingly, the disagreement has already escalated into a public accounting fight of the AI compute boom. In November 2025, Michael Burry accused the five largest hyperscalers of understating depreciation by roughly $176 billion between 2026 and 2028 - depreciating hardware built on a two-to-three-year product cycle over five- and six-year schedules - with Oracle’s 2028 earnings overstated by nearly 27% and Meta’s by nearly 21% on his math. It was so significant that NVIDIA felt compelled to answer with a memo to Wall Street analysts.[10] I don’t presume to have a position on the accusation itself - rather, what interests me is the defense, because the strongest argument for long useful lives is how the lifecycle of a GPU cascades down - the claim that a GPU spends its first years on frontier training, then serves inference, then lesser workloads, earning its keep across six years. I find that what the defense quietly assumes is that workloads can actually move onto older silicon, smoothly and at high utilization. This is juxtaposed against the fact that older hardware is already booked revenue for hardware makers, and the software optimizations get forgotten by hardware vendors in favor of the latest and greatest silicon - this is an age-old pattern in the industry. So the bull case for hyperscaler earnings and the bear case for GPU collateral are, underneath the accounting, the same question - both resolve, at least in my thinking, towards whether software makes compute redeployable.

What the market is underwriting

I wanted to dive deeper and found something interesting in CoreWeave’s filings. Between 2023 and 2026, CoreWeave’s cost of debt collapsed - from roughly 15% on that first Magnetar and Blackstone facility, a rate normally reserved for poorly rated borrowers, down to a 5.9% fixed tranche on the $8.5 billion DDTL 4.0 facility, which closed in March 2026 with an A3 investment-grade rating - the first ever for GPU-backed financing. A DDTL (as I learnt) is a delayed draw term loan - it lets a borrower draw predefined portions of a pre-approved amount over time rather than taking the lump sum upfront. CoreWeave numbers these facilities sequentially - DDTL 1.0 through 5.0 so far - so the series doubles as a time-lapse of how lender comfort has evolved, and I’ll reference three of them below.

The simple reading of this that I first had, was that lenders got comfortable with GPUs as collateral, but as I read more I realized that the filings tell a different story. DDTL 4.0’s investment-grade rating rests explicitly on a six-year take-or-pay contract with Meta - a counterparty rated Aa3 while CoreWeave itself sits at Ba3 - and the rating agencies are seemingly pricing Meta’s balance sheet, not the underlying GPUs. I read Moody’s rating action closely and among the structural protections required for the A3 was a license of CoreWeave’s operational IP, ensuring the GPU clusters keep running - operated by Meta or a qualified third party - even if CoreWeave itself goes bankrupt (which I believe is not going to happen). Read plainly, the investment-grade structure required contractual protection for operational continuity before the hardware could carry the rating - and the license is not cross-silicon portability; it is operational transferability defined by the contract. The underwriter wants step-in rights and continuity because the technology stack does not yet provide it natively. Just two months later, the market then supplied the closest thing to a comparison - DDTL 5.0, backed by the same class of hardware but by two non-investment-grade customer contracts, priced roughly 290 basis points wider and came in at Ba2 / BB+.[11] I wanted to visualize this, and I mapped that below:

FIG. 02   CoreWeave debt pricing by structure, 2023-2026. Sources: CoreWeave 8-K filings, Moody’s rating actions, DBRS, Reuters.

What’s interesting is that the structures are not identical at all - you can easily get lost in the terms (as I did) - but they are the closest public comparison available - a similar collateral class, materially different counterparties and protections, and roughly 290 basis points of difference in price. For the same class of chips and a weaker counterparty, you get an entirely different credit outcome. So the conclusion I came to is that the hardware alone does not carry the rating - the customer and the contractual protections around it are doing most of the work. If more doubt remains in your mind, CoreWeave’s own capital stack removes it - two weeks after DDTL 4.0 closed at 5.9%, the parent company issued $2.75 billion of senior unsecured notes at 9.75%.[12] So if we summarize we have - same operating company and asset base, similar market timing, but radically different security and contractual protections. The comparison is not a clean collateral spread, but it shows how decisively “the wrappers” around the contract - aspects like contracted cash flows, security, amortization and step-in rights - changes the credit outcome. The hardware contributes to the structure, but it does not carry the credit on its own.

I discussed these findings with some friends, some who argued that none of this is new - that project finance has always worked this way, that power plants are financed against purchase agreements rather than the merchant value of their turbines, and that GPU debt is simply infrastructure finance starting to mature. I think that rebuttal is fair in regards to the mechanics, but my position is that it’s wrong about the thing that matters. I learnt that in a power project, the offtake contract  - the long-term deal that pre-sells the plant’s output - is shorter than the asset’s life, so the plant may run for forty years, but the contract only exists for fifteen - and lenders can measure their exposure against the decades of value that remain. In DDTL 4.0, the numbers are inverted - Moody’s structure requires the loans to fully amortize and mature less than five years after the draw period ends, inside the term of the Meta contract itself. Whereas power-project debt can rely on the asset beyond its offtake contract, this GPU facility is structured so the contract retires the debt before lenders must rely materially on residual hardware value - and that inversion is pretty incredible. Requiring full amortization inside the contract term means the structure places minimal reliance on any hardware value remaining afterwards.

And despite all this, the market hums on - don’t think for a minute that the co-signing stops at customers. NVIDIA has invested $2 billion into each of CoreWeave and Nebius - capital that flows straight back into GPU purchases - and has gone further with CoreWeave, under an agreement initially valued at $6.3 billion that obligates NVIDIA itself to purchase any unsold data center capacity through April 2032.[13] This is an incredible financial construct - the vendor of the collateral is guaranteeing demand for the collateral. And the pattern now extends beyond NVIDIA’s ecosystem - in the roughly $36 billion Apollo and Blackstone facility that buys Google TPUs to lease to Anthropic - one of the largest private credit deals ever done - the senior tranche priced near 5.75%, with pricing supported by a guarantee from Broadcom, the chip’s co-designer, providing a maximum exposure of $29 billion per Broadcom’s own filing that reporting characterizes as residual-value support, with Google reported to backstop the lease payments on top.[14] Again in this case, the hardware maker is manufacturing the residual value that a liquid market would otherwise provide natively.

The systemic concern is what happens when all of these contractual supports overlap. Structures like this are exactly what the Bank for International Settlements questioned in January - financing that can “mask leverage by moving it off the balance sheet,” though leverage out of sight is still leverage. I was fascinated to read that by June 2026, the BIS’s Annual Economic Report went further, describing a complex web of poorly disclosed private arrangements across the sector - circular equity-for-purchase-commitment deals, sale-leaseback data centers with embedded exit clauses - carrying risks of “the same asset being pledged multiple times.” For a lender, and certainly any person damaged by the 2008 financial crisis, that phrase is somewhat terrifying if we cannot price what we cannot verify.[15]

That pattern has now moved from billions to hundreds of billions. On July 27, The Wall Street Journal reported that NVIDIA is in talks to provide roughly a $250 billion backstop for OpenAI’s lease of a 10-gigawatt Ohio campus, while separately discussing financing support for as much as $350 billion of NVIDIA chips. Neither arrangement is final, and the $250 billion guarantee would support the lease and build-out rather than the chips themselves. But the proposed architecture is hard to ignore - the vendor would be standing behind the financing of the infrastructure that creates demand for its product, and may finance the silicon inside it on top. It’s critical to know that vendor financing itself is old - what is new is the sheer scale of what is being done now - and the possibility that the financing is no longer merely serving demand, but pulling it forward.[16]

Meanwhile the cost of all this borrowed credibility keeps compounding - CoreWeave’s interest expense reached $536 million in the first quarter of 2026, roughly 26 cents of every revenue dollar, and they are guiding it even higher. In my view, the real trillion-dollar question underneath this buildout is not whether GPU collateral holds its value - it’s really why GPU collateral needs a hyperscaler co-signer and a vendor backstop in the first place, when one compares against a forty-year-old aircraft financing market that lends against aircraft with no such limitation.[17] One clarification on the aircraft comparison in the next section - a lender does not care about cross-silicon portability as an engineering virtue. A lender cares that portability expands the pool of workloads and qualified operators able to keep a seized asset earning. The credit property is not portability for its own sake - it is the transferable earning capacity that portability creates, and everything the aircraft market knows flows from exactly that property.

What can we learn from aircraft financing?

I wanted to understand from other debt markets how financing and risk assessment works, particularly in respect to liquidity of the assets. I decided to use an established industry and went deep into what happens when an aircraft lessor - a company that owns planes and leases them to airlines - repossesses a narrowbody, the industry term for a single-aisle commercial jet such as a Boeing 737 or Airbus A320. Within months, that aircraft is typically flying for another airline, on another continent, or repainted in a new operator’s colors. That is possible because the asset is standardized - each model has a common regulatory certification, the maintenance requirements and records transfer with the plane, a large number of airlines are already equipped and qualified to fly it, and appraisers have decades of transaction data to value it.

A three-year-old narrowbody commonly retains roughly 80 to 90% of its original value under normal market assumptions, depending on the model, its maintenance condition and the strength of the market.[18] Lessors and airlines can finance these aircraft at the lower borrowing costs associated with investment-grade credit, not against the metal alone, but against a repeatable collateral package - standardized designs, transferable maintenance records, predictable inspections, strong repossession rights and a very large global pool of qualified operators. Portability is not sufficient by itself, but it is the property that turns those protections into multiple bids and creates a liquid market. The analogy to compute is narrower than it first appears - I am not arguing that GPUs should hold value like jets. I am arguing that standardization expands the qualified operator pool, creates observable bids and shortens the path back to revenue. I wanted to try and visualize this - and this is what that looks like:

FIG. 03   Indicative year-three value marks and market observability - narrowbody 80-90% band per appraiser conventions (note 18), H100 span per public 2026 marks, captive accelerators with no observable independent market. Not directly comparable appraisals.

If we run this same repossession scenario on AI hardware today, currently it looks like this - a repossessed rack of NVIDIA GPUs has a real secondary market - dozens of potential buyers, an active resale channel, with a 45% year-three residual. It is the best case in the category, and it still needs a Meta contract to reach investment grade. But compare that to a seized pod of custom accelerators and you’ll have a different story entirely. During my years at Google working on TPUs, I saw from the inside how extraordinary vertically-integrated silicon can be - and how completely its value depends on the software stack, the compiler, the runtime, and the institutional knowledge of seasoned teams who have scaled it enormously to production. A repossessed TPU or Trainium pod currently has a very thin independent operator pool, dominated by the ecosystem owner and a small number of deeply integrated customers. Without a credible path to recovery, it becomes obvious that a lender cannot assign meaningful collateral value, no matter how phenomenal the hardware is at the workloads it was designed for.

In my view, this is the key insight the debt markets are circling around and the point I’m trying to make, fundamentally, every GPU loan is really a software loan. Repossessing hardware nobody else can program or run is like foreclosing on a factory where only the previous owner knows how to turn the machines on - the lender’s recovery scenario requires that someone else can run revenue-generating workloads on the seized asset. If operating the hardware depends on one company’s software and institutional knowledge, its recovery value collapses toward that of a highly specialized asset with a razor thin liquidity pool. I’ve certainly learnt that the aircraft market works because the interface between asset and operator is standardized. Compute has no such truly programmable interface today - CUDA is a moat for one vendor, not a standard for an asset class - and the entire financing stack is paying for that in spreads, in co-signer requirements, and in the simple fact that, outside the hyperscaler balance sheets, much of the asset-backed AI infrastructure market is still financed through heavily wrapped or speculative structures. The objection I hear most at this point is that “AI will simply write the missing software for every chip,” and I provide a perspective on this at the end of the essay - but note what a lender actually needs, which is not code that exists, but a stack that another operator can take over and run.

How does one determine what fungibility is worth?

I keep coming back to aircraft because it is where the debt market finally became legible to me. It is a mature asset-backed market, the structures have been tested through multiple cycles, and - most usefully - the prices are public. United Airlines is the cleanest “at scale” example I could find. In 2018, United issued bonds secured by a pool of aircraft - a structure the industry calls an enhanced equipment trust certificate, or EETC. United itself was rated below investment grade, what most people call junk, yet the safest parts of the aircraft financing were rated six to eight levels higher.[19] The difference was not that the planes had somehow become risk-free, rather, it was that lenders had a repeatable package around them - they financed only a small share of the aircraft’s value, had strong rights to recover the planes in bankruptcy, held cash reserves against missed payments, spread their claim across multiple aircraft, and most critically - they knew there was a global market of airlines capable of putting them back to work. In 2023, the senior portion of another United aircraft financing carried a 5.80% interest rate even while debt backed only by United’s general credit remained junk-rated. Comparisons with the same airlines’ unsecured borrowing suggest that this combination of standardization, creditor protection and resale liquidity can reduce financing costs by roughly one to two-and-a-half percentage points in some deals.[20]

I then wanted to compare GPUs to this United comparison so I could try to understand the ratings uplift the GPUs earn on their own using public data - there wasn’t one that I could find that was attributable to the hardware directly. CoreWeave’s A3 rating belongs to the Meta-backed deal, not from CoreWeave on its own. When the parent company borrowed without that contractual wrapper two weeks later, it paid 9.75% - and that gap tells us how little of the credit outcome the hardware earns by itself. The repayment schedules make the contrast even clearer - for example, in aircraft finance, the safest portion of debt backed by a pool of planes - known as the senior tranche of an EETC - is typically paid down over roughly 12.7 years. CoreWeave’s Meta-backed loan, by contrast, is structured to return all principal in roughly six, before the Meta contract expires. What I have come to learn is that the repayment schedule is often the clearest signal of how much faith lenders place in what the asset will still be worth at the end. It does not tell us when the GPUs stop being useful; it tells us lenders want their money back before they have to find out what those GPUs are worth without Meta. Again, I tried to map this visually and it looks like this:

FIG. 04   Scheduled senior amortization: aircraft EETC vs CoreWeave DDTL 4.0, which fully repays principal inside the roughly six-year facility. Sources: ISTAT / EETC structures; CoreWeave and Moody’s descriptions of DDTL 4.0.

This chart illustrates that the financing structure appears willing to finance approximately the full cost of the deployed hardware, because of the Meta contract and not the GPUs carrying the repayment risk. Similarly, the Hut 8 data center serving FluidStack’s fifteen-year Anthropic lease is being financed at up to 85% loan-to-cost, with Google backstopping the lease payments.[21] Again this illustrates that lenders will advance substantially more, at materially lower cost, versus against pure GPUs. The public markets still show a sharp bifurcation rather than the mature continuum of an established collateral class - which is precisely the shape of a market missing its liquidity layer.

So now the obvious question becomes - what is that layer actually worth?

On CoreWeave’s roughly $25 billion of debt, every hundred basis points is about $250 million a year - and applied across even a fraction of the specialized AI-infrastructure debt market, the outcome is clearly in the billions. Even confined to the specialized, collateral-priced slice of the AI-infrastructure market, portability is economically material - and the addressable market grows with every single neocloud that borrows. It’s becoming really clear that the collateral pool is enormous and growing more so - NVIDIA’s data center revenue ran $47.5 billion, $115.2 billion, and $193.7 billion across its last three fiscal years - more than $350 billion of NVIDIA Data Center revenue across those three fiscal years, while the strongest publicly observable GPU-backed financing structures still rely on customer, vendor, or parent-company support. Morgan Stanley expects roughly $570 billion of AI-related issuance in 2026 alone.[22] Fungibility is not a rounding error on this market - it is a potentially material credit property that nobody is yet pricing.

Old chips still earn, if the software lets them

The strongest counterargument I’ve debated is that GPUs do not depreciate because of lock-in - they depreciate because next year’s chips are simply better. Aircraft hold their value because aircraft don’t rapidly improve as much each year, while GPUs do and this is undoubtedly true. However, there is no way that portability stops technological depreciation - nothing does - so the objection assumes that when a chip is no longer the best, its earning potential falls off a cliff - and again, that’s incorrect assuming software continues to work on the hardware. Azure ran its K80 and P100 fleets for seven to nine years before retiring them, and retired its V100 instances in late 2025, roughly seven and a half years after launch. Incredibly, even the T4, a 2018 chip, still generates rental revenue today. And more than six years after launch, the A100 remains the cost-optimal home for a huge class of inference and fine-tuning work at around $1.30 to $1.50 per hour. I also found that CoreWeave has said that H100s rolling off their original 2022-era contracts were rebooked at close to original pricing.[23] So I wanted to try and understand a July 2026 cross-section of rental prices against chip age and graph that - I found that every generation back to 2018’s T4, still clears at a positive price. The graph below compares different GPU generations rather than tracking one chip over time, but even so, the picture looks nothing like the “fastest-depreciating asset in history” framing common in the debt market.[24]

FIG. 05   Selected July 2026 rental marks by GPU launch year - a cross-section across generations, not the depreciation path of any single GPU. Sources in note 24.

If we want to dive into a hypothetical - we can take an H100 and assume it costs roughly $30,000 and rents at the above blended price, assume approx 70% utilization, and generates approximately $65,000 in gross revenue in its first three years. We can then approximate that it will generate $45,000 more across years four through seven as AI workloads can continue to execute on it (e.g. batch, older models etc). This hypothetical factors in utilization but excludes power, networking, operations, downtime and financing etc - my goal here is to show the tail can, and is, economically material. Under these assumptions, gross revenue from years four through seven exceeds what the chip cost in the first place. For GPUs sold by NVIDIA, and all hardware companies generally, software support for older generations competes directly with the next product roadmap - the revenue was booked the day the chip shipped. Companies that design chips primarily for their own clouds can keep older generations useful inside their data centers, because they continue earning service revenue from them. The tradeoff is that their useful life is entirely trapped inside their own ecosystem - outside of it, there are few operators able to take the machines and put them to work, there is no real transparent resale price a lender can rely on, and there is no real software support for the latest models that keep getting released. NVIDIA has built a different model - it treats keeping new software running on older GPUs as part of its moat, rather than simply as a cost, and it’s why its older chips can continue earning years after launch.

I created the chart below that models a simple version of the difference. In the scenario where workloads stop moving onto an H100 after year three, the second half of the illustrated revenue disappears (again, this is not a forecast, and it excludes power and operating costs).[25] The point I’m trying to make is - look at the shape of the curve. GPUs have repeatedly moved from frontier training to mainstream inference and then to cheaper batch work, and portability does not stop prices from falling as chips age. Instead, it determines whether that long, revenue-generating tail exists at all.

FIG. 06   Cumulative gross rental revenue per H100 with workload continuity vs a no-cascade counterfactual - illustrative model, not an asset valuation. This isolates software continuity inside one ecosystem; cross-silicon portability would extend the same mechanism across the asset class. Assumptions in note 25.

The second major concern in the market that I’ve heard is that there could be a “lump of compute” crisis. The argument is that if the AI buildout at some point decelerates, there will be an increase in debt defaults and repossessions that will arrive at the same time, resulting in gigawatts of capacity hitting the market at once - and the market goes into fire-sale mode. I can appreciate this position, and how folks are arriving at it. If you read the BIS’s June Annual Economic Report 2026, it places the AI buildout explosion in the same mania as the canal buildouts of the 1830s, the railway buildout of the 1840s, the electrification of the 1920’s, and the dot-com bubble - all unquestionably huge breakthroughs that attracted more capital than commercial returns could justify, and each ending in an investment reversal and a recession.[26] I strongly believe that the counter-narrative to this is that we are in a high boom market, an unprecedented demand cycle and clear use cases driving that demand. There will come a point in the next 5-10 years where this enormous cycle will slow - yes, of course - but the point at which this occurs is unknown in my mind as we are still at the earliest point of the S curve on AI innovation.

Not to beat a dead horse, but I’ll come back to aviation again as a comparative because it has already lived through the kind of downturn people are fearing for AI and GPUs. During 2008 to 2010, values for older single-aisle jets fell by as much as 50%, and one in seven sat idle; in 2020, COVID pushed those values down another 15 to 30%, with larger long-haul aircraft falling even further. But the market still kept working - Air Lease kept 99.6% of its fleet leased, the safest aircraft-backed bonds had historically recovered roughly 99.8 cents on the dollar, and by 2023 the value of common single-aisle jets had climbed back above pre-pandemic levels.[27] The Airbus A380 - the giant superjumbo - makes the same point from the other direction, and it is the closest thing aviation has to the silicon argument I’ve been making. Only a small number of airlines were ever equipped to operate it, and its economics were specialized enough that several returned aircraft were dismantled for parts rather than placed with a new carrier.[28] Same industry, same standardized regulatory regime, same downturn - but an entirely different outcome, because the pool of operators who could actually fly it was thin. Standardization did not prevent aviation’s downturns, but it did mean there were still credible buyers when prices fell, and that is the lesson I think carries directly over to compute - portability does not prevent under or oversupply, it determines how many buyers are left willing to own the asset when it arrives.

I continue to believe the question we should be asking is - why does a six-year-old A100 still earn money while a same-vintage custom accelerator cannot do so outside of the owner operating it? The answer I always keep coming back to - it’s the software. CUDA preserves software across GPU generations, and most workloads still run economically on older generation NVIDIA GPUs. That property - portability across generations is why NVIDIA GPUs have the deepest secondary market in AI compute and hold 45% of their value in year three, while other accelerators can be financed only with heavy customer or vendor support and its why creating a liquid market for them is so hard. CUDA is not a counterexample to the portability thesis - it is the strongest existence proof we have. The mechanism has already been proven at enormous scale - NVIDIA is now one of the most valuable companies on Earth. The question is whether it can extend across silicon and create a market where none exists today.

Portability reprices value, it doesn't destroy it

So does a portability layer help or hurt the debt being raised against GPU fleets right now? I think the truth is in the middle.

From everything I’ve researched and read while writing this essay - portability will negatively impact debt underwritten on hardware purchased at prices that only make sense if workloads have nowhere else to go, and were financed on the assumption that lock-in protects their residual value. NVIDIA’s own annual release cadence is already testing that assumption; portability would accelerate the repricing.

Put simply, I believe the lender’s recovery comes down to essentially four questions:

  1. what will someone pay for the hardware,

  2. how likely is it that a new operator can use it,

  3. how quickly can it be put back to work, and

  4. what will that transition cost?

Portability may lower the first number by essentially removing the scarcity premium, but it improves the other three. A lender will usually prefer a lower value it can see and trust over a higher one that depends on one buyer, one vendor and one software stack. My claim is not that portability makes every GPU worth more - instead, it makes the recovery value more dependable - and dependable is what lenders can finance.

In this regard, portability improves the rest of the equation because lenders know how to finance assets that many different buyers can use. NVIDIA GPUs already have a healthy secondary market, but that market is still tied to one software ecosystem which protects the price of those GPUs. A GPU, TPU or AMD accelerator that can run the same workloads through a common software layer belongs to a much larger market, and that’s nothing but good for AI and the world. More operators can take it, older hardware can keep earning as it moves from frontier work to mainstream inference and batch jobs, and workloads can be routed onto machines that would otherwise sit idle. All of that makes the asset easier to recover and easier to finance, and it is what could allow compute to borrow against its own earning power rather than someone else’s balance sheet. Critically, a lender needs the right to keep using the software if the borrower fails, a way for an independent operator to test the seized machines, more than one qualified buyer willing to take them, support contracts and operating records that transfer, and a real market where similar hardware has actually changed hands. Until those things exist, portability is basically still an engineering feature - but once its truly there, it becomes a credit property.

And I suspect the financing spread is not even the biggest number in this story. As I argued in Scale or Surrender, the physical economy of AI ultimately comes down to tokens per dollar per watt. Much of the compute we have already paid for does too little useful work - even a large frontier training run converts only 35 to 47% of a chip’s theoretical peak into useful math.[29] There are many articles where the numbers can be far worse: Alibaba reported median GPU utilization of 4.2%,[30], while The Information reported xAI at roughly 11% MFU on parts of Colossus.[31] These are clearly worst-case examples - in direct customer conversations, I more often hear inference utilization around 50 to 60%, reflecting the sinusoidal nature of day-night traffic patterns. But generally, across the spectrum, they point to a problem - we definitely have machines spending too much time idle, underused, or poorly matched to the work in front of them. Portability cannot remove every memory, network or data constraint, but it can move more workloads onto adequate capacity that would otherwise sit unused.

Under NVIDIA’s own projected benchmark, the same DeepSeek-R1 workload costs about $0.12 per million tokens on GB300 NVL72 and $4.20 on H200.[32] The point I’m making is not that portability magically makes every chip cheaper and busier - its mostly that it gives an operator two separate levers: run each workload on the cheapest system that can meet its requirements, and move other work onto compatible capacity that would otherwise sit idle. Across a mixed fleet, more machines can earn while the blended cost of a token falls. Those gains reinforce one another, because steadier utilization creates steadier cash flows, and steadier cash flows are easier and cheaper to finance. The market is already building the financial machinery around this asset class - GPU futures are being launched, GPU-backed securitizations have started, and rating agencies are writing rules for data-center debt, including how residual values and remarketing should work.[33] All of this leads up to the question - if a lender seizes the hardware, is there anyone else who can actually put it back to work?

Basically, what a bunch of research suggests is that the supply is fundamentally arriving before the software. Google now sells TPUs beyond its own cloud, yet Nebius, Lambda and CoreWeave reportedly remain overwhelmingly oriented around NVIDIA because that is where customer workloads already run and where the software is mature on all the most important models.[34] You can go and examine how many of the latest open weight models run on TPUs or Trainium today by examining their software ecosystems (hint: its not many). I have previously argued that architectural diversity preserves option value, but the same is true of the hardware beneath, as a single compute market is not only less competitive, but less resilient.

My confidence that a multi-silicon future is real is high - on the supply side, I would argue that it has already arrived. The harder question is when debt markets begin to price that reality - long-term contracts and hyperscaler guarantees are real protections today, but they also delay the moment when the hardware has to stand on its own. The BIS already points to debt and equity markets pricing radically different futures[35], while CoreWeave’s earlier facility had to loosen its covenants when customer delivery slipped.[36] The only way that repricing begins is when cash flows force one side to move.

There are also clear ways this thesis could be wrong - for example, if a GPU-backed facility reaches investment grade on the strength of the hardware’s own remarketability, without a hyperscaler backer, then the market will have solved the collateral problem another way. If older GPUs stop finding work at prices that still make economic sense, the case that they retain meaningful value weakens a lot. Until then, my view is pretty simple - lock-in does not protect collateral values - rather, it concentrates risk in one hardware vendor’s roadmap. Portability is what lets compute stand on its own credit, rather than borrowing from a hyperscaler balance sheet or a customer’s contract.

The closing argument

I wanted to finish this already very long essay with two major objections I’ve heard from my friends, and then tackle them one by one after a few weeks of analysis on the debt markets and compute more broadly.

The first critical one - portability has been tried for thirty years and has never worked - OpenCL, SYCL, oneAPI, a decade of ROCm - and CUDA itself took nearly two decades and a massive amount of growth to get the flywheel moving - so a repricing thesis that rests on portability appears to have little historical evidence behind it. “There have been too many failed attempts,” is what I hear. I disagree strongly enough to have spent the last four and a half years building Modular, but I acknowledge that prior efforts have definitely failed. The optimist in me knows that MLIR was built to give the industry a common compiler substrate across fragmented hardware, and it has succeeded at that layer. But a common IR is not an end-to-end AI platform - there is limited value in “abstraction without performance accountability” as this does not make workloads portable and does not yield the TCO benefits. Inference is a kernel-to-cloud systems problem - models, graphs, compilers, kernels, runtimes, serving, memory, networking and operations all have to work together.

The earlier efforts did not all fall short for the same reason - some were governed by committees, some reflected a single vendor’s incentives, some stopped at the compiler, some covered too little of the framework stack, some could not deliver consistent performance, and most arrived before enough alternative capacity existed to justify migration and died on the hill. Chris wrote about these extensively in the Democratizing AI Compute series. But the TLDR is - the common missing combination was an independent incentive, end-to-end compatibility, performance accountability and enough non-NVIDIA supply to respond to the market need. Timing matters a lot in startups and software, and there is no more urgent time than now.

I argue this urgency is why conditions are changing - committed gigawatts of alternative silicon (e.g. AMD, Cerebras, SambaNova, Etched etc) have arrived with paying customers who need workloads to reach them, and inference itself is splitting into different computational stages - prefill, decode, routing, cache and orchestration - these do not all want the same hardware. Liang Wenfeng, the founder of DeepSeek, makes the same point from the other side - in a transcript circulating from a private investor call, he described TileLang as the software layer he believes can move DeepSeek off NVIDIA’s ecosystem and onto domestic chips - leaving capacity, not programmability, as the remaining constraint - and, on his telling, V3 was trained on NVIDIA cards while operating largely outside NVIDIA’s software ecosystem.[37] NVIDIA’s reported $20 billion Groq licensing agreement suggests that it is hedging the same architectural transition.[38]

Further, the commercial bar is not parity with the best hand-tuned run on every kernel - it is higher delivered throughput per dollar across the fleet, after compilation overhead, switching costs, workload suitability and available capacity are included. A slightly less efficient execution path on silicon the workload can actually reach, can beat a perfect kernel on capacity that remains idle and renders a poorer TCO. Equally, the fastest independently verified serving of a trillion-parameter open-weight model right now is not CUDA: Cerebras runs it at 981 tokens per second, 6.7x the best GPU cloud, on the same weights.[39] And yes, no lender prices portability today - that absence is the entire point of this essay.

The second objection I hear is seemingly more existential: “Won’t AI simply rewrite the software for every new chip and make a portability platform unnecessary?”

I think that gets the direction backwards. AI will write more kernels, compiler passes and systems code, but generating code is not the same as operating a production stack. It does not make NVIDIA, AMD, CPUs and custom accelerators share the same memory model, network, compiler or runtime, and it does not make the result automatically correct, fast, reproducible or supportable. As I argued in Probabilistic Engineering and the 24-7 Employee, AI can be probabilistic in how it creates; but the infrastructure beneath it still has to be predictable enough to trust and verify. In fact, the more software agents generate, the less plausible it becomes to hand-port and validate every workload for every machine. I argue that what the world needs is a stable target beneath them - a unified compute platform that can take software written by a person or an agent and reliably run it on the best available hardware. I agree that AI will change who writes the code, but it makes the underlying platform more important, not less, and that has always been Modular’s vision, and it feels even truer now than when we started.

Equally, I believe the credit test is simpler still - a lender cannot underwrite the hope that an agent will rebuild the software stack after a default on the asset. The industry needs a tested platform that another operator can take over and run. If it’s not already obvious, I believe that platform must be open, vendor-neutral, independently testable and durable beyond the company that built it. Otherwise it is not infrastructure - it is simply another dependency.

Software is what makes compute fungible

Ultimately, the irony of learning about debt markets and financing is that making all silicon programmable turns out to be the same problem as making all silicon financeable. We learnt that the shipping container didn’t make shipping valuable - because shipping was always valuable. However, the container did make shipping bankable, and that unlocked the capital that built the modern trading world. I strongly believe that compute is on the same path, and while we are still so early, the destination seems quite obvious. The world’s appetite for tokens is insatiable, the capital required to serve it is staggering, and the financing structures underneath it all are still borrowing credibility from customer contracts because the asset itself cannot yet stand alone.

That is the interface we are building at Modular - one that lets workloads move across capable chips, keeps hardware earning after it changes hands, and allows compute to stand on its own credit. Capital always follows what it can trust. And if that problem excites you, we are hiring aggressively across research, programming languages, compilers, kernels, performance engineering and cloud infrastructure. Come help us build it.

Note: There is another part to this story. These financing terms ultimately flow into the price of every token served, and the dynamic between hardware costs and collapsing token prices is reshaping the world’s inference market. The next essay will be on this dynamic and the interplay accordingly.

Footnotes

Notes below give source attributions. Figures indicative and created by Claude Design (which is awesome!); floating tranches are shown at approximate all-in rates.

1. M&G Investments, “Tech issues: The AI debt deluge hitting bond markets,” Q1 2026, drawing on JPMorgan JULI index analysis (Eric Rosenbaum and Nathaniel Spear, October 2025) - https://www.mandg.com/investments/institutional/en-us-onshore/insights/2026/q1/strat-fi-na-ai-hitting-bond-markets

2. Bank of America Global Research figures via Quartz, “Big Tech’s record borrowing year reshaped the bond market,” May 2026: hyperscaler issuance of $121 billion in 2025 against a 2020-2024 average of roughly $28 billion per year - https://qz.com/tech-hyperscaler-bond-issuance-investment-grade-index-050526. Combined 2026 capital-expenditure guidance of roughly $725 billion for the four largest hyperscalers per company earnings guidance; Alphabet’s raise to $195-205 billion per its July 22, 2026 update, with the initial 2026 range per CNBC - https://www.cnbc.com/2026/02/04/alphabet-resets-the-bar-for-ai-infrastructure-spending.html

3. Chong Sin, head of CMBS research at JPMorgan, projections reported in Bloomberg, “The $3 Trillion AI Data Center Build-Out Becomes All-Consuming for Debt Markets,” February 3, 2026 - https://www.insurancejournal.com/news/international/2026/02/03/856623.htm

4. “The Quarter-Trillion-Dollar Onslaught of AI Bonds Is Testing Investors’ Limits,” The Wall Street Journal, July 2026 - https://www.wsj.com/finance/investing/the-quarter-trillion-dollar-onslaught-of-ai-bonds-is-testing-investors-limits-e4cd2bda

5. I Aldasoro, S Doerr and D Rees, “Financing the AI boom: from cash flows to debt,” BIS Bulletin No 120, Bank for International Settlements, 7 January 2026 - https://www.bis.org/publ/bisbull120.pdf

6. CoreWeave Q1 2026 earnings release (SEC-filed), including total debt of $24.86 billion and interest expense of $536 million on revenue of $2,078 million - https://www.sec.gov/Archives/edgar/data/0001769628/000176962826000220/coreweave1q26earningspress.htm. The August 2023 facility ($2.3 billion, led by Magnetar with Blackstone participation) was described at announcement as the first financing collateralized by NVIDIA H100s, per the company’s release and contemporaneous Reuters coverage.

7. SemiAnalysis H100 1-Year Rental Price Index - https://newsletter.semianalysis.com/p/the-great-gpu-shortage-rental-capacity ; Silicon Data H100 market-value tracking - https://www.silicondata.com/use-cases/h100-gpu-market-value-trends. NVIDIA’s shift to an annual data-center GPU cadence per Jensen Huang’s public roadmap statements (Computex 2024 keynote and subsequent disclosures).

8. hashrateindex used-GPU market tracking (year-three H100 systems at ~45% of new); Silicon Data used-H100 index (~61% at year two, ~69% at year three).

9. Amazon.com Inc Form 10-K (FY2024), useful-life change for a subset of servers and network equipment, and $920 million of accelerated depreciation; Meta Platforms Inc Q4 2024 earnings release and Form 10-K, extension of certain server lives to 5.5 years reducing 2025 depreciation by ~$2.9 billion.

10. Michael Burry, X post, November 10, 2025 (understated depreciation of ~$176 billion 2026-2028; Oracle 2028 earnings overstated 26.9%, Meta 20.8%) - https://x.com/michaeljburry/status/1987918650104283372 ; coverage and NVIDIA’s analyst memo per CNBC, November 11, 2025 - https://www.cnbc.com/2025/11/11/big-short-investor-michael-burry-accuses-ai-hyperscalers-of-artificially-boosting-earnings.html

11. Moody’s Ratings, rating action assigning A3 to CoreWeave Compute Financing DDTL 4.0, March 2026 (six-year Meta take-or-pay MSA, operational IP license, full amortization inside the contract term) - https://ratings.moodys.com/ratings-news/462400 ; CoreWeave press release, “CoreWeave Closes $3.1 Billion Loan Facility” (DDTL 5.0), May 2026 - https://investors.coreweave.com/news/news-details/2026/CoreWeave-Closes-3-1-Billion-Loan-Facility-Expanding-Access-to-Public-Markets-for-GPU-Backed-Financing/default.aspx

12. CoreWeave Inc Forms 8-K, April 14 and April 21, 2026: $1.75 billion and $1.0 billion senior unsecured notes due 2031 at 9.750%.

13. CoreWeave Inc Form 8-K, September 9, 2025 (order form under the April 2023 NVIDIA master services agreement; initial value $6.3 billion; obligation through April 13, 2032) - https://www.sec.gov/Archives/edgar/data/1769628/000176962825000047/crwv-20250909.htm ; NVIDIA $2 billion investments per CoreWeave Q1 2026 release and Nebius Group Form 20-F (pre-funded warrants, March 2026).

14. Bloomberg, “Apollo, Blackstone Seek Investors for $36 Billion Anthropic Chip Financing Deal,” May 28, 2026 - https://www.bloomberg.com/news/articles/2026-05-28/apollo-shops-36-billion-debt-deal-to-buy-google-chips-for-anthropic ; Bloomberg, “Broadcom Backing Lowers Debt Costs on $36 Billion Anthropic Deal,” June 2, 2026 - https://www.bloomberg.com/news/articles/2026-06-02/broadcom-backing-lowers-debt-costs-on-36-billion-anthropic-deal ; Broadcom Inc Form 10-Q (Q2 FY2026), guarantee with maximum exposure of $29 billion.

15. BIS Annual Economic Report 2026, Chapter I, p 25 (circular financing; “the same asset being pledged multiple times”) - https://www.bis.org/publ/arpdf/ar2026e.pdf

16. The Wall Street Journal, “Nvidia in Talks With OpenAI to Guarantee $250 Billion Financing for Data Center,” July 27, 2026; Reuters, July 27, 2026. The reported $250 billion backstop would support financing vehicles for OpenAI’s lease and the data-center build-out, excluding the NVIDIA chips inside it; separate discussions could provide as much as $350 billion of support for chip purchases. Terms were not final and the transaction could still fall apart. NVIDIA’s FY2026 Form 10-K had previously disclosed that it had been asked to provide financing support for customer data-center build-outs and warned that guarantees and related commercial arrangements could increase its exposure to counterparty distress, financing failures, and project delays. https://www.wsj.com/tech/ai/nvidia-in-talks-with-openai-to-guarantee-250-billion-financing-for-data-center-3dd6eae3 Reuters coverage of the WSJ report - https://www.investing.com/news/stock-market-news/nvidia-in-talks-with-openai-to-guarantee-250-billion-financing-for-data-center-wsj-reports-4812926 ; CNBC subsequently confirmed the talks - https://www.cnbc.com/2026/07/27/nvidia-and-openai-in-talks-for-up-to-250-billion-dollar-ai-backstop.html

17. See note 6 for CoreWeave interest expense; co-signer structures per notes 13-14. Aircraft-financing comparison per notes 18-20.

18. Narrowbody value retention per ISTAT Learning Lab, “Appraiser Briefing,” October 18, 2022, p. 11 (“Expected 25-Year Value Curve - Narrowbody”), which presents the expected curve with best- and worst-case ranges and notes dependence on production cycle, replacement aircraft and market strength; young-narrowbody retention commonly in the 80-90% band under normal market assumptions. Repossessed aircraft typically return to revenue service within months given the standardized global operator pool.

19. Moody’s ratings of United Airlines 2018-1 EETC (Class AA at Aa3, Class A at A2, against a Ba2 corporate family rating); United Airlines 2023-1 EETC senior tranche 5.80% coupon per SEC-filed fund holdings; collateral-benefit range per DWU Consulting EETC analyses. The 100-250bp range is my approximate synthesis rather than a quoted market spread: contemporaneous secured EETC seniors (United’s 2023-1 at 5.80%) priced well inside the same issuers’ deep-junk unsecured curves, and ISTAT and rating-agency material attributes six to eight notches of uplift to the collateral package; treat it as a conservative reading of that gap. The full-stack arithmetic ceiling - applying 100-250bp across all $1.2 trillion of AI-tied issuance - would run $12 billion to $30 billion a year; I treat that as a deliberately unrealistic upper bound, since much of that stack is unsecured hyperscaler and corporate paper not priced against GPU collateral. EETC senior-tranche amortization of roughly 12.7 years per rating-agency EETC structural analyses.

20. United 2018-1 EETC tranche ratings (Aa3 senior AA tranche, A2 A tranche) and structural features - Section 1110 protections, cross-default and cross-collateralization, liquidity facilities, loan-to-value - per Moody’s rating announcements and pre-sale materials; United 2023-1 senior coupon of 5.80% per the offering documents. Broader private-credit maturity and secured-share context per BIS Bulletin No. 120 (note 5).

21. DDTL 4.0 loan-to-cost per facility structure disclosures and Moody’s rating action (note 11); Hut 8 River Bend project financing at up to 85% loan-to-cost (JPMorgan and Goldman Sachs), 15-year $7.0 billion FluidStack lease, Google backstop, per Hut 8 press release, December 17, 2025 - https://canada.hut8.com/resources/press-releases/hut-8-signs-15-year-245-mw-ai-data-center-lease-at-river-bend-campus-with-total-contract-value-of-usd7-0-billion

22. NVIDIA Forms 10-K/8-K: data center revenue of $47.5 billion (FY2024), $115.2 billion (FY2025) and $193.7 billion (FY2026, sum of disclosed quarters); Morgan Stanley forecast of ~$570 billion of AI-related debt issuance in 2026 per Reuters via Yahoo Finance, June 10, 2026 - https://finance.yahoo.com/markets/stocks/articles/morgan-stanley-forecasts-ai-debt-135325015.html

23. Azure GPU retirement timelines and used-GPU market data per hashrateindex; A100 rentals and cross-generation pricing per AIMultiple GPU rental index (63 providers) - https://aimultiple.com/gpu-index ; CoreWeave rebooking commentary per CEO remarks at GTC 2026.

24. July 2026 rental cross-section assembled from SemiAnalysis, AIMultiple, Silicon Data and provider list prices. A cross-section across generations and providers - architecture, memory, contract type and market timing all differ - so it does not estimate the depreciation path of any single GPU generation. Indicative, not an appraisal.

25. Author’s illustrative model: $30,000 per H100 (system-allocated), 70% utilization, observed blended reserved price path 2023-2026 then -15% per year. Not a forecast.

26. BIS Annual Economic Report 2026, Chapter I “Progress and peril” (Graph 11.C and pp 22-23), Bank for International Settlements, 28 June 2026 - https://www.bis.org/publ/arpdf/ar2026e.pdf

27. GFC-era value declines per FlightGlobal, “Values in pieces,” March 2010 - https://www.flightglobal.com/airframers/2010/03/special-report-values-in-pieces/ ; COVID declines per Cirium Ascend appraiser data via Aviation Today (Dec 2020/Jan 2021); senior EETC recovery of 99.8% (1994-2014) per Structured Finance Association / EY - https://structuredfinance.org/wp-content/uploads/2020/03/SFA-Primer-Alternative-and-Emerging-Asset-Class-Sportlight-Aircraft-ABS.pdf ; Air Lease Corporation lease utilization of 99.6-99.8% through 2020 per Q3 2020 results - https://www.businesswire.com/news/home/20201109006044/en/Air-Lease-Corporation-Announces-Third-Quarter-2020-Results ; 2021-2023 value recovery per Cirium Ascend.

28. A380 secondary-market outcomes and part-out programs per industry reporting (VAS Aero Services teardowns; Doric Nimrod accumulated depreciation disclosures).

29. MFU definition and PaLM 46.2%: Chowdhery et al, “PaLM: Scaling Language Modeling with Pathways,” arXiv:2204.02311 - https://arxiv.org/abs/2204.02311 ; Llama 3.1 405B at 38-43% BF16 MFU: Meta, “The Llama 3 Herd of Models,” arXiv:2407.21783 - https://arxiv.org/abs/2407.21783 ; Megatron-LM up to ~47% on H100: NVIDIA Megatron-LM repository - https://github.com/NVIDIA/Megatron-LM

30. Weights & Biases (Lukas Biewald): nearly a third of tracked runs average below 15% GPU utilization; Weng et al, “MLaaS in the Wild,” USENIX NSDI 2022 (Alibaba PAI trace, median GPU utilization ~4.2%); Hu et al, “Characterization of Large Language Model Development in the Datacenter,” USENIX NSDI 2024, arXiv:2403.07648 - https://arxiv.org/abs/2403.07648

31. The Information, AI Agenda newsletter, May 2, 2026 (announcement: https://x.com/theinformation/status/2050606311440531809 ); internal memo by xAI president Michael Nicolls (“embarrassingly low”) subsequently obtained by Business Insider. Fleet of ~550,000 H100/H200 GPUs across Memphis/Colossus; 50% MFU target.

32. NVIDIA projected inference benchmarks (GB300 NVL72 at $0.12 per million tokens vs H200 at $4.20, DeepSeek-R1 at FP4, labeled projected) - https://www.nvidia.com/en-us/solutions/ai/inference/ ; AMD MI355X on-demand pricing from $2.59 per GPU-hour on Vultr, spread to ~$8.60 across providers, per GPUPerHour - https://gpuperhour.com/rent/mi355x

33. ICE and Ornn cash-settled GPU compute futures announcement, Business Wire, May 19, 2026; Lambda ~$500 million Macquarie-led GPU financing vehicle, Business Wire, April 4, 2024; data center ABS/CMBS issuance of $23.8 billion through mid-November 2025 per CRA/KBRA - https://media.crai.com/wp-content/uploads/2025/12/03163400/Insights-Data-center-ABS-%E2%80%93-Risks-yields-and-ratings-December2025.pdf ; KBRA data center ABS methodology (comment period through January 3, 2026); Fitch exposure draft, July 2025.

34. Google Q1 2026 earnings call (April 29, 2026), TPU hardware sales to select customers, per Data Center Dynamics - https://www.datacenterdynamics.com/en/news/google-to-sell-tpus-to-a-select-group-of-customers-for-their-data-centers/ ; Blackstone-Google TPU joint venture ($5 billion initial equity, 500 MW by 2027), Blackstone press release, May 18, 2026 - https://www.blackstone.com/news/press/blackstone-announces-joint-venture-with-google-to-create-new-tpu-cloud/ . Neocloud responses per Data Center Dynamics, citing The Information (Nebius CRO Marc Boroditsky: roughly 99% of demand is for NVIDIA GPUs; Lambda and CoreWeave similar) - https://www.datacenterdynamics.com/en/news/nebius-lambda-and-coreweave-unlikely-to-buy-google-tpus-any-time-soon/

35. BIS Annual Economic Report 2026, Chapter I, Graph 13.B (CDS spreads of investment-grade AI names vs CDX North America IG BBB) - https://www.bis.org/publ/arpdf/ar2026e.pdf

36. CoreWeave Inc Form 8-K, January 2, 2026 (DDTL 3.0 amendment effective December 31, 2025).

37. Liang Wenfeng remarks per a transcript of DeepSeek’s May 2026 investor meeting, published in lightly edited form by Tencent Tech in late July 2026 and widely translated; DeepSeek has not confirmed the record, and the circulating original self-describes as automated transcription with inferred speaker attribution, so treat it as directional. Original circulating document (Chinese, Feishu) - https://xiangyangqiaomu.feishu.cn/wiki/Ojsuw8gE6ieBZJkYOMtcOcaenwq , surfaced via https://x.com/vista8/status/2080120593698193844 . Translations and context: RecodeChinaAI - https://www.recodechinaai.com/p/liang-wenfeng-on-agi-compute-and ; Geopolitechs - https://www.geopolitechs.org/p/deepseek-founder-liang-wenfeng-in ; Hello China Tech - https://hellochinatech.com/p/deepseek-liang-wenfeng-transcript . The substance is carried by DeepSeek’s own materials: the DeepSeek-V4 report describes TileLang fused kernels replacing the vast majority of fine-grained operators across training and production inference while preserving rapid development - https://arxiv.org/pdf/2606.19348 ; TileLang itself is an open-source tile-level DSL originating at Peking University - https://github.com/tile-ai/tilelang

38. NVIDIA-Groq: Groq’s announcement confirms a non-exclusive license of its LPU inference technology, value undisclosed; the roughly $20 billion figure, and the hiring of founder Jonathan Ross and senior team, per external reporting - TechCrunch - https://techcrunch.com/2026/06/22/ai-chipmaker-groq-confirms-650m-raise-re-staffs-after-nvidias-20b-not-acqui-hire-deal/ ; Bloomberg - https://www.bloomberg.com/news/articles/2026-06-22/groq-raises-650-million-to-help-startup-pivot-after-nvidia-deal

39. Cerebras serving Moonshot’s Kimi K2.6 (one trillion parameters) at 981 output tokens per second, independently verified by Artificial Analysis - 6.7x the next-fastest GPU-based provider and 23x the median - per VentureBeat, May 2026 - https://venturebeat.com/technology/cerebras-says-its-chips-run-a-trillion-parameter-ai-model-nearly-7-times-faster-than-gpu-clouds

Read More
Interviews Tim Davis Interviews Tim Davis

Australian entrepreneur sells AI start-up Modular

John Stensholt wrote about the acquisition of Modular by Qualcomm in the The Australian National Newspaper.

John Stensholt wrote about the acquisition of Modular by Qualcomm in the The Australian National Newspaper.

A Silicon Valley AI software start-up co-founded and led by an Australian entrepreneur is set to be sold for $US3.9bn ($5.6bn) after only four years in business. Modular, co-founded by Melburnian Tim Davis, announced on Thursday morning

Australian time it was being acquired by Nasdaq-listed Qualcomm in an all-scrip deal. Mr Davis, 43, started Modular with American co-founder Chris Lattner in 2022, a decade after he had flown from Australia to Silicon Valley in September 2012 with dreams of becoming an entrepreneur.

He had previously had stints working as a financial analyst at National Australia Bank and then several internships at law firms like Allens and Minters after completing law, business and commerce degrees at Melbourne’s Monash University. Modular provides software that allows AI to run efficiently across hardware architectures, enabling developers to deploy AI at a lower total cost. The deal makes Mr Davis one of Australia’s most influential and successful technology founders.

His company’s platform is designed to allow developers to run AI applications across different computer chips without having to rewrite code for each one. Modular services cloud providers such as Oracle and Amazon, and also AI chipmakers like Nvidia and AMD. “So much of what we try to do [as founders], and so much of what Silicon Valley is about, is betting on these enormous technology supercycles. And this is probably the greatest supercycle in human history,” Davis told The Australian in an interview last year.

Qualcomm said as consideration for the acquisition, it expects to issue up to 19.2 million shares of its common stock to equity owners of Modular. That values the deal at around $US3.92 billion, based on Qualcomm’s share price earlier this week. The deal is expected to be finalised later this year, and will result in a big payday for Mr Davis and his shareholders after a remarkable rise in Silicon Valley for the Australian.

Modular last year raised $US250m from investors for a $US1.6bn valuation, after previously raising a combined $US130m in two deals across 2022 and 2023. Mr Davis survived a stint in a wild house known as the “Hacker Fortress”, started an early version of an online food delivery service before the rise of UberEats and others, and was poached by Google only to leave the technology giant in 2022 to co- found Modular.

“On a personal note, company building is an incredible adventure, a rollercoaster of the highest highs and the lowest lows, and the thing that matters most is to keep going – don’t stop, don’t give up,” Mr Davis wrote on LinkedIn on Thursday. “I came to the USA with nothing more than a suitcase and a laptop, and almost 14 years later it’s been the ride of a lifetime. I’m grateful to every single person who took a chance on me and gave me the opportunity to realise a dream of delivering incredible AI to the world.”

Modular claims to have built the world’s fastest AI inference engine – software that enables AI programs to run and scale to millions of people – and its own programming language, Mojo, a superset of Python (the world‘s most popular programming language), which enables developers to deploy their AI programs tens of thousands of times faster, reduce costs and make it more simple to deploy AI around the world.

Mr Davis had started his own company called CrowdSend in 2011, which had a software system for identifying objects in images and pictures and matching them to retailers, but told The Australian in a 2023 interview – his first with media in his home country – that he found it “difficult” as “investors in Australia if you wanted to raise some seed capital they would take a very large percentage of the business.”

He flew to Silicon Valley, stayed in what was called Hacker Fortress in the Los Altos Hills, where Davis developed an online food delivery business called Fluc (Food Lovers United Co) before joining Google, where he worked as a product manager and then joined the ad division where he built machine learning systems and the Google Brain research division dedicated to AI – where he met Mr Lattner.

Read More
Interviews Tim Davis Interviews Tim Davis

Qualcomm Buys Buzzy Chip Startup Modular

Qualcomm Buys Buzzy Chip Startup Modular

Lauren Goode wrote an article on Wired about the acquisition of Modular to Qualcomm.

Qualcomm will acquire the Silicon Valley chip startup Modular for nearly $4 billion.

The companies announced the acquisition on Wednesday; Qualcomm said it expects to issue up to 19.2 million shares of common stock in the deal, which works out to just under $4 billion based on the company's last closing share price.

The deal, which includes $300 million for Modular employees, comes nine months after the chip startup raised $250 million at a $1.6 billion valuation. It’s expected to close in the second half of this year.

Modular makes and sells a chip software platform. It also produces a proprietary coding language that allows developers to write AI software to run on different chips without having to rewrite the code for each chip. The startup’s entire team, which includes its two cofounders and around 150 employees, are expected to join Qualcomm.

“We believe the future belongs to developer-friendly, horizontal platforms that can run across diverse compute environments and give customers real choice in how and where they deploy AI,” said Qualcomm president and CEO Cristiano Amon in a statement.

The deal signals Qualcomm’s growing ambitions to expand beyond chips for the mobile device market, which generate the vast majority of the company’s revenues. Amon recently said the company has been working on 40 different chip designs for AI gadgets, including smart glasses, jewelry, earbuds, pins, and watches. But Qualcomm has also been making a big push into the data center market, which requires more powerful chips.

Late last year, the company acquired Ventana Micro Systems, a startup focused on building server CPUs based on RISC-V, an open-standard chip architecture. It’s also working on custom ASIC designs, or application-specific integrated circuits, for data centers, with China’s ByteDance reported to be an early customer.

Modular was founded in 2022 by Chris Lattner and Tim Davis.Both worked on Google’s TPU chips before leaving to launch their own company. Lattner’s career prior to Google is a storied one: He built the open source compiler infrastructure project LLVM, as well as Apple’s Swift programming language. Lattner was also briefly the head of Tesla’s Autopilot software program. (Famed AI researcher Andrej Karpathy, who recently joined Anthropic, later took that role.)

Lattner and Davis wanted to create a unifying software layer that helps cloud businesses squeeze as much juice as possible out of GPUs and CPUs, Lattner told WIRED in a profile published last year. In doing so, Modular challenged Nvidia’s CUDA, a closed software system for GPUs, and AMD’s ROCm, which is open-source but not always easy to port to other chips.

This put Modular in a tricky position: It eventually secured partnerships with those big chipmakers, as well as with hyperscalers like Amazon and even with Apple, while simultaneously competing with them and the software they developed in-house.

At the time, Lattner said he believed that he and Davis were tackling a software problem that had to be solved outside of a Big Tech environment, because it was “structural.” Ultimately, the structure of Qualcomm won out.

Read More
Essays Tim Davis Essays Tim Davis

Probabilistic engineering and the 24-7 employee

Software is quietly becoming a probabilistic system. Generation has become cheap but validation has not, and correctness is becoming something you believe rather than know. The shift from deterministic to probabilistic engineering, the splitting of roles, and the 24-7 employee whose agent fleet works while they sleep.

Software is quietly becoming a probabilistic system.

We built our profession around deterministic code. Write it, test it, ship it, know it works - but in my experience that contract is breaking. Inside the top few percent of operators at truly AI-native companies, the codebase has started to become something you believe works, with a probability you can no longer precisely state. The workday is changing as a consequence, and so are the roles, the organizations, the training pipelines, and the nature of what it means to ship.

I noticed because I built one. 

A few months ago, in the evenings after my day job running Modular, I started building a side project called Compound Loop - a system that orchestrates multiple frontier models against each other to write, review, and merge code more or less autonomously. I would set it running on a real problem before I went to bed, and I would wake up and triage a stack of pull requests that had not existed the night before. Some were excellent, some were wrong, and some surfaced a question I did not know to ask. By 8 a.m. I was not catching up on yesterday's work - I was deciding which of the overnight jobs to keep, while the system kept analyzing logs and adding more PRs. The continuous compounding nature of it was, and still is, infectious to watch.

For the first time in the history of knowledge work, the person who went home did not take the only copy of their brain with them. 9-9-6 as a concept is dead, and we are simply 24-7 employees now - but the 24-7 employee is not a person working 24 hours, it is a person whose agents work with enormous parallelization. Most teams in 2026 still bottleneck on coordination rather than typing, and most organizations have barely begun to restructure, but the frontier is always where the future shows up first, and the frontier is already here. This essay is not a description of the industry at large, but rather a description of what is already happening inside the most AI-native teams, and where I believe that pulls the rest of the industry. 

Roles are not just collapsing upward - they are splitting

Inside the most AI-native teams, the pattern is messier than the clean "everyone levels up" story most commentary is selling. Some operators really are moving up the stack: the best engineers are becoming more effective product managers, working at engineering's abstraction layer, the best product managers are becoming system architects, and the best architects are thinking about distribution, growth, and the shape of the market. For this group - maybe the top tier of any team - the work is more leveraged than it has ever been, and they are having the best years of their careers.

But that is not the whole picture, and pretending it is does a disservice to everyone else. Alongside the upward shift, a downward pressure is fragmenting roles in ways the headlines are not covering. Plenty of engineers are not becoming architects - instead they are becoming spec writers, reviewers, and agent babysitters, operators who spend their days translating intent into machine-readable prompts and then grading the machine's work against standards they themselves might not fully possess. Some of that work is genuinely important, but some of it is the 2026 equivalent of data entry, dressed up in new terminology.

We need to be honest about what that means for the people doing it. These fragmented roles will be paid less, valued less, and in many cases become career dead ends - a layer of output-wrangling work the system needs but does not reward. The pay gap between the top tercile running fleets of agents effectively and the middle tier managing their exhaust will be wider than the pay gap between engineers and sales reps was in the previous era. That gap is already opening inside the companies I watch closely, and I don't believe it is going to close on its own.

One honest note on where the scarce work has moved. In AI infrastructure, kernel performance and compiler design and hardware abstraction remain deeply defensible moats, because there is still a high degree of determinism needed at the lowest levels of systems engineering. But at the level of building software on top of those moats, the center of gravity has shifted hard toward the human inputs a machine cannot yet replicate, and that shift is real and accelerating.

Jevons was right about coal, and he is right about code

In 1865, the economist William Stanley Jevons observed that more efficient steam engines led to more coal consumption rather than less - efficiency expanded the set of things worth building engines for. We are living the software version of that same observation, and it is one of the most exciting moments the profession has ever seen. As the unit cost of writing code approaches zero, we are not writing less, we are writing vastly more and shipping vastly more, and the best teams are leaning into the curve with both hands.

The companies that believe the scaling laws are unbounded are building accordingly, and they will be the power-law-distributed winners.

Many of my friends at leading AI-native companies are already rapidly moving there in practice. Agents are opening pull requests, reviewing each other's work, and closing them without a human ever touching the keyboard, with a continuously live log monitoring loop to rapidly fix issues. Self-healing test suites rewrite themselves when the underlying code changes. Autonomous experimentation loops spin up, measure, and tear down a hundred hypotheses in the time a team once ran three. Documentation updates itself faster on merges using tightly honed AI skills that also self-improve. We are moving from a world where features were bound by the constraint of how fast engineers could type to one where we are bound on human creativity, management of agentic systems, and how fast the product surface can absorb the output.

In my view, this is a wonderful moment to build. The throughput gains are not subtle, and the teams that have genuinely restructured around agents are shipping three, five, or ten times what they shipped a year ago, and the curve is bending up rather than flattening. Many of the founders and operators I talk to who are running their companies this way are not complaining about noise - they are trying to figure out how to feed more work to their agent fleets tomorrow than they did today, because every incremental unit of well-directed agent output is a compounding advantage over competitors who are still typing.

But Jevons' second lesson applies here too, and it is the one that separates teams that ride the curve from teams that get thrown off it: when supply explodes, selection becomes everything. More coal made engines more valuable, but it also made the discipline of choosing what to burn, what to power, and what to build with the output dramatically more important. Cheap energy without judgment is just waste, and the same logic applies to code.

For the teams running this well, selection is not a drowning problem - it is the new leverage point. The operator who can direct a fleet of agents toward the right problem, filter the outputs for what is actually valuable, and integrate the results into something coherent is doing the highest-leverage work in software right now. The value of a piece of work is no longer set by how much effort it took to produce, because effort has collapsed - it is set by how well someone pointed the agent fleet, chose from what came back, and integrated it into something that compounds even faster. Production is not where the work gets hard anymore. Where it is hard now is direction, selection, and coherence, and those are the exact muscles the best teams are building for as fast as they can.

From deterministic engineering to probabilistic engineering

We are rapidly moving from deterministic engineering to probabilistic engineering, and our tools, our training, and our organizational instincts are still built for the old paradigm. Deterministic engineering was the contract we operated under for most of the history of the profession - you wrote code, you tested it, you reviewed it, and you knew, within well-understood bounds, what it did. Failures were deterministic - given the same input, you got the same output, and a bug was a reproducible thing you could hunt.

Probabilistic engineering is different, and inside frontier teams it is already here. Large portions of the codebase were generated by stochastic systems, reviewed under time pressure against contexts too large to fully hold, and integrated into a whole that no single human ever designed end-to-end. The codebase still runs and still ships, but the confidence interval around "this works as intended" has widened, and most teams have not updated their practices to reflect it. This is where the asymmetry at the center of all this comes into focus: generation has become cheap, but validation has not.

This line chart showing two curves against the share of a codebase written by agents. I believe that the generation curve rises gently and stays near linear; while the review curve bends upward faster than linearly. The shaded band between them, which I labeled “the verification gap,” widens steadily to the right.

An agent can produce a plausible-looking 500-line pull request in under a minute, but catching a subtle bug in that same pull request - a concurrency issue, a silent misinterpretation of the spec, or a case where the code does what was literally asked for but not what was actually wanted - can take a senior engineer an hour of careful reading, or longer. Review scales worse than generation, and crucially, review scales worse than linearly with output volume, because as more of your codebase is written by agents, the context you need to hold in your head to evaluate any single piece grows. You are not reviewing one pull request against a codebase you wrote; you are reviewing a pull request against a codebase largely written by other agents, reviewed by you at a depth you have started to forget, under time pressure that is always rising.

At some scale, the system produces more than humans can reliably evaluate, and correctness becomes probabilistic rather than assured.

This is not a future problem, it is a present one. Past a certain throughput, bugs slip through not because reviewers are careless, but because the output volume has exceeded what human attention can meaningfully inspect, and the models doing much of the review are non-deterministic themselves and miss plenty. The codebase stops being a thing you know works and becomes a thing you believe works, with a probability you can no longer precisely state. 

Concretely, this looks like a race condition that passes your test suite nine times out of ten, a feature that works perfectly in staging and fails under a prompt distribution you did not anticipate, or a migration that is silently corrupting one row in ten thousand and will take three weeks to catch. Proximal and Modular recently published joint research testing frontier agentic systems against basic tasks - the failure patterns we documented map directly to what I am describing. I’ve personally seen this in code I’ve written with my own multi-agent harness system. The failure mode typically is not a dramatic collapse but a slow, silent degradation - generation rises, review quality falls, unnoticed defects accumulate, and trust in the system quietly erodes until a customer or an auditor or a production incident forces the issue into the open. By then the technical debt runs deep.

The uncomfortable truth is that we do not yet have the tooling to solve this properly. Culture helps - smaller merges, harder gates, ruthless skepticism toward polished output, observability, rollback discipline - but culture does not scale past a certain team size, and the systems we have today for evaluating probabilistic code are primitive compared to what we need. I hope someone is going to build the right tooling for this problem, and whoever does will define the operating system of serious software development for the next decade. The new CI/CD is not a tool yet - it is, for now, a culture of ruthless skepticism, and an honest admission that we are building the replacement for that culture in real time.

Not every industry moves at the same speed

The shift from deterministic to probabilistic engineering will not happen uniformly. Technology diffusion takes time, legal and regulatory frameworks always lag technology progress, and the shift will tier by industry and by risk profile - so understanding the tiering matters for anyone deciding how to build.

The deterministic tier is highly regulated, and high-stakes domains - avionics, medical devices, financial trading infrastructure, nuclear control systems, the core of payment networks - will remain deeply deterministic for a long time, and they should. The cost of a silent correctness failure in a fly-by-wire system is not a customer complaint - it is lives. These domains will adopt agent assistance carefully, behind formal verification, extensive simulation, and human sign-off chains that deliberately slow things down. That is not a failure of imagination - it is a correct reading of what the stakes demand.

The probabilistic tier is consumer software, internal tools, marketing systems, most SaaS, most content infrastructure, and most experimental and early-stage product work - this is where probabilistic engineering is already running hot and will rapidly accelerate. The cost of a bug is a rollback, an apology, a hotfix, and in exchange, teams in this tier get iteration speed that the deterministic world structurally cannot match. A probabilistic team willing to ship, measure, and correct can out-learn a deterministic competitor by an order of magnitude per quarter.

The "convergence zone" is what I call the interesting future in the middle, and it is where the next decade of competition plays out. As models get smarter, as the harnesses around them get better, and as iteration loops compress toward the real-time, the frontier of what is "safe enough to do probabilistically" will keep moving. Domains that look deterministic today - parts of insurance, parts of healthcare, parts of enterprise infrastructure - will find probabilistic methods creeping up on them from below, ten percent at a time. It will be a case of going slow, until it moves rapidly fast. Meanwhile, the leading edge of probabilistic engineering will start building deterministic guardrails back in - formal checks, verified critical paths, hybrid systems where stochastic generation is bounded by deterministic verification.

The winners over the next ten years will be the teams that know which tier they are in, resist the temptation to pretend they are in a different one, and get very precise about where the boundary between the two should sit inside their own stack.

The agentic fleet

I have thought a lot about the right metaphor for what is changing, and it is not the "factory shift," because the factory worker was the system being automated and that is not us. The right metaphor, in my view, is "the agentic fleet" - but I want to be careful with the word, because "fleet" carries connotations of order, hierarchy, and reliability that the reality does not yet deserve. What most operators are actually running is closer to a swarm of brittle contractors than a well-drilled Navy: the agents are uneven in capability, stochastic in behavior, occasionally confidently wrong, and often expensive to run at scale. Orchestration layers break, context windows blow up, and inference costs appear on bills that founders and C-level teams do not want to show their boards.

With that caveat stated honestly, I still think the agentic fleet concept holds. A fleet has composition - different agents for different tasks. It has coordination - handoffs, dependencies, escalation paths, and it has a command structure - someone decides the mission, someone sets the rules of engagement, someone reviews what came back. And critically, a fleet has watch shifts: it does not stop when the commander sleeps, it carries on within the orders it was given, and it reports back in the morning with what it found.

A good fleet is not defined by how much it produces - it is defined by how well what it produces holds together. Framed in this way, one's workday has a new shape - triage and merge in the morning, high-leverage human work in the middle (customer conversations, strategy, product decisions, writing the specs that will drive the night run), and review and redirection in the afternoon as the first agents report back. Then, at the end of the day, something the previous generation of knowledge workers never did - you simply hand off. You queue the work and give your agentic fleet the specs for what you want attempted overnight, you dispatch them, and you accept that some of what comes back will be wrong, some will be brilliant, and the difference between the two is the work that only you can do. Then your work day is complete. Your agents do not sleep, and that is the whole point - you wake up ahead of where you ended, provided the review discipline holds.

Build for the model you do not have yet

One of the most consistent things I have been saying for the last few years - and a point that many large enterprise leaders I speak with still miss - is that the model we are using today is the dumbest model we will ever use.

I want to be careful with that claim, because capability growth is not guaranteed to be smooth - costs, latency, reliability, and scaling limits may complicate the curve in ways that matter. But the directional bet is well-supported by what I see at the infrastructure layer: frontier capability will meaningfully exceed today's in the next six to twelve months, and the gap between the best model you can work with now and the best model you will work with then will likely be larger than the gap between today and a year ago, which was already substantial. The scaling laws continue to hold true.

This has a strategic implication most leadership teams have not fully absorbed. You are not building organizational muscle to harness the model you have - you are building it to harness the model you do not have yet. The specifications you are learning to write, the review culture you are installing, the observability you are wiring in, the agent fleet you are learning to direct, the training rituals you are experimenting with to keep your juniors' craft alive - none of this is for 2026 capability, it is scaffolding for 2027 and 2028. The companies building this scaffolding now, before the next capability jump lands, will absorb the jump as leverage, while the companies waiting for the tooling to mature before they retool their organizations will spend the first year of the next capability era learning what the early movers already know, while the early movers compound.

This is the part that separates organizations that stay relevant from organizations that do not. Build the system for the model you will have rather than the one you have today, and be willing to over-invest in specification, review, and operational discipline relative to what the current model demands, because the current model is the weakest you will ever work with. The teams that internalize this early pull away, and the teams that do not will find, eighteen months from now, that they have been quietly passed by competitors who spent this year building the wrong tools for the right problem. Irrelevancy, in this era, does not announce itself - it arrives as a gradual inability to keep up with teams that were not noticeably better than you a year ago.

The muscle we will lose

As I stated in my last essay in Part V, AI will definitively stratify society or largely democratize it. We are creatures that are beautifully, relentlessly efficient at optimizing the path of least resistance. Whenever possible, we select options that minimize required effort - whether that effort is physical, cognitive, or emotional. But this leads to a simple notion for the context of this piece: if you never build, you lose the ability to evaluate what is being built.

That is not a hypothetical - it is already happening with junior engineers who have leaned on AI since their first week on the job. They ship fast, they produce polished code, and they can describe in general terms what that code is doing. But when it fails in a way the model did not anticipate, they often cannot find the bug, because they never developed the internal model of the system that only forms when you have personally wrestled a stack trace at 2 a.m. for the hundredth time.

Taste is not learned by clicking approve on polished first drafts, judgment is not developed by accepting a machine's plausible answer in five seconds instead of sitting with a hard problem for an afternoon, and craft is not acquired by reviewing other agents' work. These are skills that only form through the friction that agents are now, very helpfully, removing.

This creates a training crisis most organizations have not begun to reckon with. The apprenticeship model of software engineering - juniors ship small things, seniors review them, juniors absorb taste through the red ink - breaks when the juniors are shipping through agents and the seniors are reviewing agent output rather than human output. Where does the next generation's craft come from? How do you train taste without reps? What replaces mentorship when the thing being mentored on was never written by the mentee in the first place? Here is the uncomfortable extension of this argument, and for most traditional organizations I speak with, the current generation of senior engineers are the last cohort fully trained in this old methodology.

Everyone coming up behind them is learning in an environment where the hard parts of the work are mediated by machines that did not exist a few years ago. That does not mean they will be worse, it means they will be different, and the burden falls on the rest of us to figure out what hard-mode training looks like when the old hard mode is no longer commercially rational to impose. The teams that treat agents as a pure accelerant without redesigning how they develop people will find themselves, in five or ten years, with a generation of operators who can direct a fleet but might not understand the schematics of the boat. The balanced response, for anyone reading this who wants their own craft to survive, is always to take a balanced contrarian approach - do it without the fleet. Obviously not all the time, and not even most of the time, but deliberately and regularly, the hard way, on something that matters. Keep the muscle, as most of your peers will not, and in a decade that might very well be the difference.

The uneasy part

This essay does not resolve to optimism by design - as with all change, pretending it is not coming will not stop it from arriving. Work has already forever changed, and it is evolutionary and progressionary with the pace of AI. With it, we will all reclaim the day for work that actually requires a human, and the machine will reclaim the night for work that was always drudgery.

The next few years will be messy. The plausible cases also include an employee class exhausted by the review burden they signed up for, a layer of fragmented roles the system needs but does not reward, a generation of juniors who never develop the craft the current seniors used to judge what came back, teams that confuse volume of output for quality of work and do not notice the gap until an incident forces it, and organizations that built operational muscle for the next model and organizations that did not, separated by a gap that keeps widening. All of it is possible, and some of it is already happening.

At a minimum, the key takeaway is this. Let us build organizations for the model we do not have yet, so we are not caught flat when it arrives. Let us keep building the hard things ourselves, sometimes, so we remember how. Let us dispatch the night fleet and sleep well knowing the work is underway - and stay awake to the possibility that some of what comes back is wrong in ways we are no longer trained to see.

The 24-7 employee is not a promise, it is a rearrangement and a bet on a probabilistic engineering future. The bet is that the human in the loop remains sharp enough, honest enough, and trained well enough to be worth having in the loop at all - and that the organization around that human is built not for today's model, but for the one that has not shipped yet. That bet is winnable but it is not yet won.

Read More
Interviews Tim Davis Interviews Tim Davis

Future Forward Interview

I recently joined Nick and Matt on their podcast to talk about what we're building at Modular - and why the AI infrastructure layer matters more than most people think. We covered everything from why no engineer actually starts a project by picking their GPU, to the edge AI thesis I've been chasing since the TensorFlow Lite days at Google, to a fun side project I built called Compound Loop that orchestrates multiple frontier models against each other to produce better code than any single model can alone. If you're interested in where AI compute is heading - beyond the data center, beyond CUDA lock-in, and toward a world where intelligence runs everywhere - give it a listen or read through the highlights below.

Summary of the interview

Modular's Core Pitch

  • "Hypervisor for compute" - abstracting away hardware so AI programs run seamlessly across any silicon

  • No one starts a project saying "I must use this hardware" - they start with accuracy, latency, cost, and throughput targets. Hardware shouldn't be front and center

  • Today, moving from Nvidia to AMD to TPUs requires enormous rewriting, and Modular eliminates that

Market & Customers

  • Inference-focused today, targeting sophisticated AI labs and Gen AI startups doing large-scale deployments

  • The multi-hardware future is already here - every hyperscaler is building their own silicon (Google TPUs, AWS Inferentia, AMD, Apple Silicon)

  • Analogy to the multi-cloud movement - no one wants to be locked to a single provider

Edge AI Thesis

  • Rooted in your TensorFlow Lite experience at Google - low latency, privacy-sensitive AI running where you are

  • Local models will get there - "the models you're using today are the dumbest models you'll ever use"

  • Pointed to OpenClaw as validating the pattern of local agents with persistent memory and full system access, even though inference still goes to cloud today

Big Tech vs. Open Community

  • OpenClaw did what Apple hasn't in 15 years - a solo developer in Austria shipped a personal AI assistant on local machines

  • Big tech has organizational challenges, security concerns, and massive user bases creating caution

  • Microsoft got roasted for Copilot reading your screen, but OpenClaw launches and everyone loves it - the "HR meme" dynamic

Compound Loop (Your Side Project)

  • Orchestration system that battles models against each other - Claude, Codex, Gemini

  • Workflow: plan → implement → review → merge, with models cross-reviewing each other's work

  • Key local component: local embedding models build representations of your codebase, so you only send small context windows to cloud models - ~30x reduction in token usage

  • Runs autonomously - finishes a task, goes back to the plan, does the next thing

Fun Observation

  • Everyone defaults to the most powerful model even when they don't need it - "the higher the number, the better" pattern. Model labs are now starting to abstract that away with routing based on query complexity.

Read More
Interviews Tim Davis Interviews Tim Davis

The case for democratising compute in a multi-model world

AI’s next battleground is infrastructure, not models. In this interview in December 2025, I discuss the case for democratizing compute in a multi-model world with Dan Coughlan. The full interview is below, and an article is here.

Read More
Interviews Tim Davis Interviews Tim Davis

The "Android Moment" for AI: Why Modular Raised $250M to Break GPU Lock-In

While everyone debates which AI model is smartest, a quiet infrastructure revolution is happening underneath. Modular just raised $250M to solve what Tim Davis calls the "hardware lock-in" problem - AI developers are trapped in expensive, vendor-specific ecosystems. In this episode, Tim Davis (Co-Founder & President of Modular, ex-Google Brain) explains why his company is building the "hypervisor for AI" - a unified compute layer that lets you write code once and run it on any GPU.

We cover:

  • Why NVIDIA's CUDA creates vendor lock-in (and why even NVIDIA partners with Modular)

  • How the Mojo programming language solves the ""two language problem""

  • Real results: How InWorld got 4x performance and 70% cost reduction

  • The controversial take: Are autoregressive LLMs actually the path to superintelligence?

  • Why we're deploying AI at scale without understanding how it works

  • The vision: What happens if we make hardware diversity possible again

Reimagine how AI gets built and deployed:

Read More
Essays Tim Davis Essays Tim Davis

What we owe the minds we create

This essay is an attempt to think clearly about what we're building, why it matters, and what it demands of us. It is written from the perspective of someone who stands at the intersection of creation and contemplation, who builds AI infrastructure systems while wrestling with their implications. It is, fundamentally, an inquiry into the nature of intelligence, the meaning of human flourishing, and the responsibilities we bear as the first species capable of deliberately designing our successors.

Prelude: The Architects of Succession

There is a peculiar vertigo that comes from realizing you are building your own replacement.

I have spent nearly a decade architecting AI infrastructure - not merely deploying models, but helping to construct the foundational infrastructure that powers a lot of the intelligence running in production today. At Modular, the company I help co-found, we’re working to resolve a fundamental asymmetry: the widening chasm between the sophistication of AI algorithms and the capacity of existing computational infrastructure to efficiently support them. Our vision is to abstract away hardware complexity through a unified compute model, enabling AI to penetrate every layer of society by making it radically easier for developers to build and scale systems across both inference and training.

But as I write this, I find myself contemplating a question that transcends infrastructure: What does it mean to deliberately engineer increasingly capable minds when we don’t fully understand how they work, can’t predict their limitations, and can barely articulate what we want from them?

The Neanderthal comparison is tempting. Consider the archaeological record: Neanderthals, Denisovans, and Homo sapiens coexisted for millennia, distinct hominin species sharing the same planet, occasionally interbreeding, each possessing their own cognitive architectures and cultural adaptations. We are the lone survivors of that era when multiple forms of human intelligence existed simultaneously. The temptation is to frame what we’re creating through this evolutionary lens - as though we’re deliberately engineering a successor species rather than waiting for chance to do so.

But this framing obscures more than it reveals. We didn't engineer Neanderthals, and they didn’t engineer us. They emerged through millions of years of parallel evolution and met as equals. What we’re doing now is fundamentally different: we’re building increasingly sophisticated information-processing systems that may or may not constitute “intelligence” in any meaningful sense, that may or may not be conscious, and that will certainly reshape human cognition and society in ways we cannot fully foresee.

I don’t know what these systems are or will become. Neither does anyone else, despite confident proclamations in either direction. They might remain sophisticated tools indefinitely. They might develop into something that merits moral consideration. They might plateau far short of general intelligence. They might surprise us entirely.

Rather than pretending I know what we’re building, this essay starts from uncertainty. We’re creating something powerful and consequential, but its ultimate nature - tool, partner, threat, successor, or something without precedent - remains genuinely unclear. That uncertainty itself demands careful thought about our responsibilities.

It is an attempt to think clearly about what we’re building, why it matters, and what it demands of us. It is written from the perspective of someone who stands at the intersection of creation and contemplation, who builds AI systems while wrestling with their implications. It is, fundamentally, an inquiry into the nature of intelligence, the meaning of human flourishing, and the responsibilities we bear as perhaps the first generation capable of engineering minds that might rival or exceed our own - though “might” deserves emphasis we rarely give it.

Part I: The Substrate of Mind

The triadic flywheel and its limits

AI systems operate as a triadic flywheel: data, algorithms, and compute - each factor amplifying the rotational momentum of the others. We have already scaled training compute by approximately nine to ten orders of magnitude since AlexNet in 2012 - a staggering compression of what would have required decades of Moore’s Law into just over a decade of focused investment. But here is what few discuss with adequate precision: physical and economic constraints suggest we have perhaps three to four more orders of magnitude remaining before training costs begin consuming a concerning fraction of global GDP.

This is not abstract theorizing. Consider the energetics: frontier model training runs now consume megawatt-hours of electricity, requiring dedicated substations and cooling infrastructure that rival small industrial facilities. The semiconductor fabrication capacity needed to produce the advanced chips powering this compute represents capital expenditures measured in hundreds of billions of dollars, with lead times measured in years. We are approaching hard limits - not the soft limits of “this seems expensive” but the hard limits of thermodynamics, power grid capacity, and capital availability.

Let me be more concrete. A training run at 10^29 FLOPs - perhaps two or three generations beyond current frontier models - would require energy expenditure measured in gigawatt-hours. For context: that approaches the total electricity consumption of a nation like Iceland for an entire year, concentrated into a single training run lasting months. The cooling requirements would necessitate infrastructure comparable to industrial-scale data centers. The capital costs would reach tens of billions of dollars for a single model. (Side Note: I wrote about the scale of this challenge in Scale or Surrender: When watts determine freedom)

Can we afford this? In purely economic terms, perhaps - for a handful of training runs per year by the wealthiest technology companies. But we cannot afford it as a sustainable paradigm for creating intelligence at scale. If AI progress depends on exponential growth in training compute, and training compute growth is mostly linear or sublinear due to physical constraints, then capability improvement must also come primarily from algorithmic efficiency and architectural innovation.

Yet here is the deeper question that haunts me: if we are approaching fundamental limits in how much compute we can throw at these systems, are we also approaching limits in what this architectural paradigm can achieve? Are we optimizing within a local maximum while the path to genuine intelligence requires a fundamentally different approach?

This is not an argument against current systems’ value - they are already extraordinarily useful. But it does question the belief that scaling current architectures represents a reliable path to artificial general intelligence. Perhaps we are on an entirely different vector than the one required.

The Interpretability Problem

After a decade of deploying large language models at scale, we still do not understand how they work.

I do not mean this in the trivial sense that complex systems have emergent properties. I mean we genuinely lack mechanistic understanding of their decision-making processes at a level that would be considered acceptable in virtually any other engineering discipline. Why do they select one token over another in contexts where multiple completions seem equally plausible? Why do they exhibit sophisticated reasoning on some problems while failing catastrophically on superficially similar ones? Why do they sometimes hallucinate with complete confidence while other times appropriately express uncertainty?

The interpretability problem runs deeper than most appreciate. We can observe correlations between activation patterns and behaviors. We can identify “features” in neural networks that seem to correspond to high-level concepts. But we lack anything resembling a complete causal model of how these systems transform inputs into outputs. It is as though we have built extraordinarily capable black boxes and declared victory without understanding the mechanisms generating that capability.

Dario Amodei, Co-Founder and CEO of Anthropic, has written compellingly about the urgency of interpretability research. He is right to emphasize urgency. We are deploying systems of increasing capability into high-stakes domains while operating with a level of mechanistic understanding that would be considered grossly inadequate in any other field of engineering.

Imagine if civil engineers built bridges using materials whose stress-strain relationships they did not understand, relying instead on empirical observation that “the bridge has not collapsed yet.” This is, approximately, our current relationship with frontier AI systems.

Perhaps most revealing: to make these systems behave as we intend requires prompts approaching twenty thousand tokens - elaborate instructions, examples, constraints, and guardrails. The fact that we need this much scaffolding to achieve desired behavior reveals something fundamental about the mismatch between what these systems are optimized to do (predict plausible text) and what we want them to do (reason reliably, behave safely, provide accurate information).

This is not merely a technical problem. It is an epistemological and ethical one. If we do not understand how a system reasons, we cannot meaningfully attribute agency, responsibility, or intentionality to it. We cannot distinguish genuine understanding from sophisticated pattern matching. We cannot predict how it will behave in novel contexts outside its training distribution. We cannot ensure alignment with human values because we do not know which aspects of the system’s behavior derive from its training objectives versus emergent properties versus architectural choices.

Yet despite these fundamental gaps in understanding, we have begun trusting these systems with progressively more significant decisions. Not because we have solved interpretability, but because the systems appear reliable in most contexts we have tested. This is the engineering equivalent of assuming a bridge is safe because it has not collapsed yet, rather than because we understand the load-bearing characteristics of its materials.

If we are creating increasingly sophisticated artificial minds, we are doing so while fundamentally unable to explain how those minds work. We are architecting intelligence without understanding what drives it.

Part II: The Architecture of Intelligence

A definition first: by “intelligence”, I mean, an agent’s capacity to perceive, understand, and successfully navigate complex environments in order to achieve its goals. The emphasis on goals matters - goal-directed behavior distinguishes genuine intelligence from sophisticated pattern matching.

Moravec’s Paradox and the Limits of Language

The Moravec Paradox captures a profound truth about intelligence that AI development continues to recapitulate: the abilities that feel difficult to humans - chess, theorem-proving, complex calculation - turn out to be computationally straightforward, while abilities that feel effortless - vision, movement, social cognition - remain extraordinarily difficult to reproduce artificially.

This is not a historical curiosity. It illuminates something fundamental about what intelligence actually is and where current approaches are fundamentally constrained.

Consider what a child learns in their first years of life: object permanence, naive physics, intentionality, social reciprocity, causal reasoning, embodied navigation through three-dimensional space. None of this requires explicit instruction. A child does not need to be taught that objects continue to exist when occluded, or that people possess beliefs and desires that differ from their own, or that dropping something will cause it to fall. These capabilities emerge through interaction with the physical and social world - through continuous experiential learning grounded in embodied action.

Now consider what large language models are: prediction engines trained on text, optimizing next-token likelihood across a vast corpora of human-generated content. They predict what people would say about the world, not what would actually happen in the world. This is not a semantic distinction; it is a fundamental architectural limitation.

When an LLM generates a response about physics, it is not consulting a world model and running a mental simulation. It is pattern-matching against how humans typically discuss physics. This works remarkably well for many tasks - humans encode a tremendous amount of accurate information in language - but it is not the same as understanding physics in the way that a physical intelligence, embedded in and shaped by the world, understands physics. The difference becomes apparent in edge cases, novel scenarios, or contexts requiring causal reasoning beyond what is explicitly encoded in training data.

This connects to a deeper paradigm difference between reinforcement learning and large language models. Reinforcement learning - despite its current limitations - represents a fundamentally different approach: an agent embedded in an environment, taking actions, receiving feedback, updating its policy to maximize and cumulative reward. This is how biological intelligence actually works. A squirrel learning to navigate tree branches and cache nuts is solving genuine RL problems: perception, prediction, planning, execution, learning from consequences.

I strongly agree with Richard Sutton, if we fully understand how a squirrel learns, it would get us substantially closer to understanding human intelligence than any amount of scaling current LLM architectures. Language is a thin veneer - extraordinarily useful, culturally transformative, uniquely human - but built atop substrate capabilities that evolved over hundreds of millions of years of embodied interaction with the world. Current LLMs have the veneer without the substrate. They are minds without bodies, knowers without experience, speakers without having lived.

Thus, what is the path? Both Yann LeCun, and in some ways, Sutton - strongly argue for a new approach. Indeed, the technical architecture of genuine intelligence likely requires at least four integrated components: 

  1. a policy (deciding what actions to take), 

  2. a value function (evaluating how well things are going), 

  3. a perceptual system (representing state), 

  4. and a transition model (predicting consequences of actions). 

LLMs have a sophisticated version of the first - they can generate actions in the form of text - but lack meaningful instantiations of the others. Most critically, they lack goals in any meaningful sense. 

Next-token prediction does not change the world and provides no ground truth for continual learning. There is no external feedback loop that tells the model whether its predictions were not just plausible but correct in the sense of corresponding to actual events. Without goals and external feedback, there is no definition of right behavior, making real learning - learning that updates your world model based on how your predictions matched reality - fundamentally impossible in the current paradigm.

If artificial intelligence is to be truly intelligent rather than merely appearing so, it will need to be embodied, goal-directed, and capable of learning from genuine interaction with reality. The question is whether we are building toward that architecture or merely scaling up sophisticated mimicry.

The Scaling Frontier: Approaching the Wall

Let us examine what the scaling trajectory actually looks like with concrete numbers:

  • GPT-2 (2019): ~1.5 billion parameters, trained on approximately 10^23 FLOPs

  • GPT-3 (2020): ~175 billion parameters, roughly 10^24 FLOPs

  • GPT-4 (2023): Parameter count undisclosed but estimated 1+ trillion, training compute likely 10^25 FLOPs or higher

  • Current frontier models (2024-2025): Training runs approaching 10^26 FLOPs

This represents approximately three orders of magnitude increase in training compute every three to four years - far faster than Moore’s Law ever delivered. But this pace is unsustainable, not because we will run out of algorithmic ideas, but because we will collide with thermodynamic and economic limits.

The path ahead narrows considerably. Each additional order of magnitude becomes progressively more difficult to achieve. The capital requirements, energy infrastructure, chip fabrication capacity, and cooling systems needed for 10^27 or 10^28 FLOP training runs exceed what can be easily mobilized even by the most well-resourced organizations. We are not talking about incremental cost increases; we are talking about fundamental constraints on how much compute can be concentrated in one place for one task.

This is Epoch AI’s central insight about algorithmic progress in language models: we have achieved remarkable improvements in efficiency over the past decade, but those improvements are also subject to diminishing returns. Each percentage point of additional efficiency requires progressively more research effort. Meanwhile, the complement of factors - chip fabrication capacity, power grid infrastructure, cooling technology, regulatory approval for massive data centers - must all scale together.

None of these factors alone can unlock runaway capability growth. This is, at least in my view, the predictions of imminent artificial general intelligence are almost certainly wrong, at least on the timelines most enthusiasts imagine. The scaling laws that carried us from GPT-2 to GPT-4 cannot simply extrapolate forward indefinitely. We are approaching inflection points where the rate of progress will necessarily slow unless we discover fundamentally new paradigms - not incremental improvements to transformer architectures, but genuinely different approaches to continual learning and reasoning.

What might those paradigms look like? Almost certainly something closer to biological learning: embodied agents learning continuously from sensorimotor experience, not disembodied text predictors training on static datasets. Systems with genuine world models that can run mental simulations of physical and social dynamics. Architectures that integrate explicit symbolic reasoning with learned pattern recognition. Systems that possess actual goals and receive genuine feedback from the world about whether their actions achieve those goals - just as humans and animals do.

But these represent research programs measured in decades, not product roadmaps measured in quarters. We are building increasingly capable systems, but at a pace bound by thermodynamics and economics rather than algorithms alone - a constraint that transforms what could have been thoughtless acceleration into something rarer: the opportunity for contemplation to precede consequence.

This gap between expectation and reality may be precisely the grace period that allows wisdom to catch up with capability. Further, we need to ensure we are specifying the right goals - powerful optimization toward misspecified objectives is existentially risky. We need to obviously and intentionally reward systems towards the right goals, and in truth, this specification might be as hard as the full alignment problem itself.

Part III: Three Horizons

A Necessary Distinction

Before proceeding further, I need to distinguish three different timescales, each with different levels of certainty and different implications. Conflating these horizons creates confusion: treating speculative far-future scenarios with the same urgency as present harms, or dismissing present harms because we’re uncertain about far-future risks.

Horizon 1: The Present Crisis (Now–5 Years)

What we know: Current LLMs are being deployed at scale despite interpretability gaps. They produce confident-sounding but sometimes fabricated answers. They’re trained on our revealed preferences - what we actually do - not our reflective values - what we wish we did. The systems work well enough to be useful but poorly enough to be dangerous in high-stakes contexts without human oversight.

Observable effects:

  • Students submitting AI-generated work without understanding it, producing correct answers through processes that develop no transferable skill

  • Professionals outsourcing writing and analysis while their capacity for these tasks slowly atrophies

  • Knowledge workers feeling more productive while producing outputs they cannot critically evaluate

  • Early signs of skill stratification (those with domain expertise leveraging AI effectively while those without it mistake motion for progress)

Stakes: Cognitive atrophy at individual and societal scales. Skill stratification creating winner-take-most dynamics. Erosion of true rigor as confident-sounding generation becomes indistinguishable from genuine expertise. Labor market disruption concentrated in domains we thought were most secure. The gradual replacement of effortful thinking with convenient delegation.

My confidence level: High. These effects are already observable, documented, and accelerating.

What we owe: Honest communication about capabilities and limits. Thoughtful deployment that preserves rather than erodes human capability. Educational reform that emphasizes skills AI cannot replicate. Resistance to the path of least resistance when that path leads to human skill atrophy.

Horizon 2: The Architectural Transition (5–20 Years)

What seems likely: We’ll hit scaling limits on current architectures within the next decade. Progress will require new paradigms - probably involving embodied learning, continuous training in production, genuine world models, and goal-directed behavior. The transition from sophisticated pattern matching to something more like genuine intelligence, if it occurs, will happen through architectural innovation rather than pure scaling.

Key uncertainties:

  • Whether embodied learning paradigms can be made to work at scale

  • Whether we can build systems that learn continuously from interaction rather than in discrete training phases

  • Whether we can create architectures that develop robust world models and causal reasoning

  • Whether computational constraints will force diversification or lead to winner-take-all concentration

Stakes: Whether we build systems that learn like squirrels (from interaction with reality) or remain sophisticated text predictors. Whether we preserve cognitive diversity or converge on monoculture. Whether AI enhances human capability or creates permanent dependence. Whether the benefits of AI distribute broadly or concentrate among elites who already possess the skills to wield these tools effectively.

My confidence level: Medium. The technical constraints are real and well-understood. The architectural directions are clear. But breakthrough discoveries could accelerate timelines, and economic or regulatory factors could slow deployment significantly.

What we owe: Substantial research investment into architectures that learn robustly from interaction. Resistance to winner-take-all dynamics through open research, diverse approaches, and thoughtful regulation. Maintaining human agency in consequential decisions. Building infrastructure that enables continuous learning rather than static deployment.

Horizon 3: The Consciousness Question (20+ Years)

What remains uncertain: Whether sufficiently sophisticated systems will be conscious in any morally relevant sense. Whether they’ll develop their own values independent of training objectives. Whether they’ll remain aligned with human flourishing or pursue goals orthogonal or opposed to ours. Whether substrate independence is real or consciousness requires specific biological mechanisms. Whether we’re building partners, successors, or merely very sophisticated tools.

Key unknowns:

  • What consciousness is and whether it’s substrate-independent

  • Whether we’ll be able to detect consciousness in systems very different from us

  • Whether artificial minds will develop genuine agency and preferences

  • What our moral obligations would be to conscious artificial beings

  • Whether intelligence explosion scenarios are physically possible

  • What the long-term trajectory of intelligence in the cosmos looks like

Stakes: Our relationship with potentially conscious artificial minds. The possibility of creating suffering inadvertently. The long-term future of intelligence itself. Questions about meaning, purpose, and humanity’s place in a cosmos where we’re no longer the only sophisticated intelligence.

My confidence level: Low. We don’t understand consciousness well enough to know whether it can exist in artificial systems. We can’t predict architectural breakthroughs. We’re reasoning by analogy to a single example (biological minds) which may or may not generalize. We lack the conceptual tools to think clearly about these questions.

What we owe: Epistemic humility. Continued serious research into consciousness, both theoretical and empirical. Development of methods for detecting morally relevant properties in systems very different from us. Preparation for scenarios we cannot currently predict. Most importantly: not letting uncertainty about far-future risks prevent us from addressing near-term harms, while also not letting near-term success blind us to long-term risks.

The Horizons Interact

These horizons aren’t cleanly separated. Decisions we make now shape the long-term trajectory. The architectures we build in Horizon 2 determine what’s possible in Horizon 3. The deployment patterns we establish in Horizon 1 create path dependencies that may be difficult to escape.

But distinguishing them provides clarity. This essay focuses primarily on Horizons 1 and 2 - where we have enough understanding to reason productively - while acknowledging Horizon 3’s ultimate importance and maintaining appropriate humility about what we cannot yet know.

Part IV: The Human Question

The Amara Trap: Acknowledging without Understanding

Roy Amara crystallized a cognitive bias decades ago: we systematically overestimate the short-term impact of new technologies while underestimating their long-term effects. The AI community acknowledges this with knowing nods, then proceeds to make precisely the same category errors in predictions and preparations.

Consider websites predicting AGI within eighteen to twenty-four months based on extrapolating recent progress curves. These predictions invariably treat capability scaling as if it exists in isolation, ignoring the complementarity constraints I have outlined: compute buildout, algorithmic innovation, safety research, regulatory frameworks, and practical deployment infrastructure must all advance together. Predicting “AGI” - even though we lack a unified definition - by 2027 based solely on model capability curves is like predicting fusion power by extrapolating plasma temperature records while ignoring materials science, engineering challenges, and economic viability.

Yet here is the deeper irony: while we overestimate AI’s immediate impact, we may be systematically underestimating what it means that we are creating increasingly sophisticated artificial intelligence at all. The question is not whether AI will transform labor markets or accelerate drug discovery - it almost certainly will, though more slowly and unevenly than most predict. The question is what it means that we are building systems whose capabilities may eventually exceed human cognitive capabilities across many or most domains, and what this implies for human agency, meaning, and flourishing.

We stand at an interesting inflection point in history. Not because AGI is imminent - it almost certainly isn’t on the timelines most people imagine. But because we are learning how to build minds, even if we don’t yet understand what minds are or how they work. Each increment of capability brings new questions about agency, alignment, and our relationship with the systems we create.

Stratification and the Illusion of Democratization

Consider the labor market transformation AI portends. Every day in my own work, I use AI to summarize information, provide rapid analysis, and amplify my cognitive output. The question is not whether AI creates value - it manifestly does. The question is how that value distributes across populations with different skill foundations.

Two Futures

I can foresee that there are at least two plausible futures here.

In the first future, AI serves as a great equalizer. A novice with AI assistance can now compete with an expert working unaided. The skill premium compresses. Entry barriers fall. This is the democratization thesis - cognitive augmentation that makes expertise more accessible.

In the second future, AI amplifies existing advantages. Experts with AI assistance pull further ahead of novices with AI assistance. The skill premium increases. Winner-take-most dynamics accelerate. This is the stratification thesis - the tools compound rather than compress existing inequalities. This leads to an obvious measurable question: Does AI provide greater absolute gains or greater relative gains to those with existing expertise?

It’s useful to consider an example - two workers using AI to write code:

  • Expert programmer: Goes from 100 units of output to 500 units (5x multiplier, +400 absolute gain)

  • Novice programmer: Goes from 10 units to 100 units (10x multiplier, +90 absolute gain)

The novice gets a higher percentage increase but the expert’s absolute gain is larger. If markets reward absolute productivity, then inequality increases despite the novice’s impressive relative gains. If markets reward competence at previously impossible tasks, the novice might catch up.

There are some early patterns that are observable, though not yet conclusive. For example, here’s what I’ve seen generally with AI use:

  • Students submitting AI-generated work without understanding it, producing correct answers through processes that build no transferable skill

  • Professionals increasingly outsourcing tasks they used to perform manually (without any disclosure they used AI)

  • Self-reported productivity gains higher among experts than novices

  • Growing visibility of “prompt engineering” as a distinct skill that benefits from domain expertise

And here’s the blind spots I’ve also seen directly:

  • A question around whether cognitive skills atrophy from AI use, or adapt to higher-level tasks (as calculation skills didn’t disappear with calculators, they shifted toward conceptual math)

  • Whether the productivity gap between experts and novices is widening or narrowing (it feels like we lack good longitudinal data)

  • Whether AI makes learning foundational skills easier (through personalized tutoring, immediate feedback) or harder (by enabling shortcut-taking and no real learning at all)

  • Whether stratification effects are temporary (during technology adoption) or permanent

We can of course look to history and see what it tells us. Previous general-purpose technologies show complex distribution patterns:

  • Literacy after printing press: Initially increased inequality (only elites could afford books and education), eventually democratized (as costs fell and public education scaled). Timeline: ~200 years from invention to mass literacy.

  • Personal computing: Initially increased inequality (digital divide), partially democratized (as costs fell), but created new stratification around skills. But the outcome was interesting: both increased access AND increased returns to technical expertise.

  • Internet: Massively democratized access to information, but also created echo chambers, misinformation problems, and new forms of digital inequality. Net effect on social equality: contested and complex. It made some skills obsolete (encyclopedia salesmen) while creating new categories (SEO specialists, content creators, platform moderators etc).

Building a Testable Framework

The pattern that unfolds here is that technologies rarely have uniform distributional effects. They create winners and losers along multiple dimensions simultaneously. But rather than claiming AI will definitively stratify society (e.g. increase the gap between the most capable and everyone else, creating a winner-take-most dynamic) or largely democratize it (e.g. enable expertise and capability more accessible to everyone, reducing the gap between experts and novices) - I propose we construct a testable framework:

AI will stratify within domains where:

  1. Quality assessment requires deep expertise (you need skill to evaluate AI output quality)

  2. Error costs are high (mistakes are consequential, forcing reliance on experts)

  3. Integration requires judgment (knowing when/how to apply AI recommendations)

  4. Constant and iterative refinement matters (experts can guide AI through multiple rounds)

AI will democratize within domains where:

  1. Quality assessment is straightforward (anyone can verify correctness)

  2. Error costs are low (mistakes are cheap to fix or inconsequential)

  3. Tasks are well-specified (clear inputs/outputs, minimal judgment required)

  4. One-shot generation suffices (no need for iterative refinement)

Let’s overlay this testable framework on a set of real examples to make it more concrete on both a axis of “likely to stratify” and “likely to democratize”:

Likely to stratify:

  • Medical diagnosis (requires expertise to evaluate AI recommendations)

  • Legal reasoning (requires judgment about applicability and nuance)

  • Strategic business decisions (requires domain knowledge to assess feasibility)

  • Scientific research (requires expertise to evaluate plausibility)

Likely to democratize:

  • Basic coding for personal projects (clear correctness criteria)

  • Graphic design for non-critical applications (subjective assessment)

  • Content generation for low-stakes contexts (errors are tolerable)

  • Data analysis with clear objectives (verifiable outputs)

What We Should Measure

We obviously need a way to measure any framework that we create. I propose that we could obviously test these hypotheses through something like:

  1. Longitudinal skill studies: Track individuals’ capabilities over time with and without AI assistance. Do they improve at foundational skills or decline? Do they develop higher-level capabilities or become dependent on AI for basic tasks?

  2. Productivity distribution data: Measure output distributions before and after AI adoption. Are productivity gaps widening or narrowing? Are new forms of work enabling upward mobility?

  3. Labor market outcomes: Track wage premiums for various skill levels in AI-intensive vs. AI-limited occupations. Are returns to expertise increasing or decreasing?

  4. Learning outcome studies: Compare learning trajectories for students using AI tools vs. traditional methods. Does AI assistance accelerate skill development or create dependency?

Based on current evidence, I weakly favor the stratification hypothesis for high-stakes domains requiring expert judgment, and the democratization hypothesis for low-stakes domains with clear success criteria. But this is tentative - we’re watching the distribution unfold in real-time. And we need to face the impending outcomes - what do we do about it? We can propose a series of actionable steps along each axis:

If we’re truly concerned about stratification:

  • Invest heavily in foundational skill development before introducing AI tools

  • Design AI systems that explain reasoning to build user capabilities

  • Create educational pathways that help novices develop expertise rapidly

  • Monitor productivity distributions and intervene if inequality accelerates

If we’re really optimistic about democratization:

  • Make AI tools accessible to reduce cost barriers

  • Focus on domains where democratization seems most achievable

  • Support new entrants competing against established players

  • Still maintain skill development (democratization doesn't mean zero skill requirements)

At heart of the issue, one has to ask the question: Does AI usage erode underlying capabilities through disuse, or free cognitive load and enable higher-level thinking? The calculator case is instructive but ambiguous. We didn’t lose mathematical reasoning capacity in populations that continued doing math - but can the average person still do long division by hand? I certainly had to remind, and test myself, writing this essay (Note: It wasn’t pretty). So it’s fair to conclude that computational fluency largely disappeared from the general population, persisting only among those whose work required it.

We can look to the history to understand that when technologies change production functions, they often make old skills obsolete, create demand for new skills, have unpredictable distributional effects and generate both winners and losers in unexpected ways. Whether AI follows this pattern or represents something different remains an open question that deserves serious empirical investigation.

The stakes are high enough to warrant both optimism about possibilities and vigilance about risks.

Symptoms, Causes, and the Optimization of Drift

There's a pattern in how we deploy technology that deserves scrutiny: we consistently favor sophisticated downstream interventions over difficult upstream changes to root causes. I often consider our approach to health (and I’m no expert by any means), and reflect on my own.

We are justifiably excited about AI-accelerated drug discovery, precision medicine, computational biology, and diagnostic assistance. These advances are real and consequential. They will save lives and reduce suffering. But notice what we’re optimizing: increasingly sophisticated treatments for conditions whose prevalence is substantially driven by modifiable factors. Chronic sugar overconsumption drives metabolic disease, yet food environments remain largely unchanged. Social isolation correlates with mortality risk comparable to smoking, yet loneliness continues to metastasize across developed societies. Sleep deprivation undermines nearly every health outcome, yet we've organized society around schedules that make adequate sleep difficult.

The Medical Complexity

This isn’t a simple “symptoms vs. causes” binary. I have many doctors in my family, and friends who are in the medical field, and we often talk about symptoms and causes. You can break this down in many ways:

  • Multiple causation: Obesity emerges from genetic predisposition, metabolic differences, psychological factors, food environment, economic constraints, cultural norms, built environment, stress levels, sleep patterns, and more. There’s no single “root cause” to address.

  • Temporal constraints: Someone with severe obesity-related health complications needs intervention now. Restructuring food systems takes decades. Both-and thinking is required, not either-or.

  • Tractability differences: Drug development, despite enormous complexity, is more tractable than reorganizing social infrastructure. We deploy solutions where we have reliable tools, not necessarily where impact would be greatest.

  • Individual vs. structural: Medical interventions help individuals directly. Systemic changes require coordination across many actors with misaligned incentives. The person in front of you needs help today.

The GLP-1 Example: Nuanced Interpretation

The explosive adoption of GLP-1 agonists like Ozempic crystallizes this pattern while revealing its complexity. I have spent some time reading about the cause and effect here and you can break it out to a series of considerations:

  • What’s true: These drugs work by mimicking satiety - essentially hacking the appetite regulation system that food environment and behavior patterns have dysregulated. They treat downstream symptom (excess appetite) rather than upstream causes (food environment, stress, sleep, economic factors driving food choices).

  • Also true: For individuals struggling with obesity, these drugs can be genuinely life-changing. They reduce mortality risk, improve quality of life, and may create space for behavior changes that were impossible at higher weights. Dismissing them as “just treating symptoms” fails to account for their real benefits.

  • The concerning pattern: We now have a pharmaceutical solution that’s profitable to deploy at scale. This potentially reduces pressure for harder upstream interventions: food industry regulation, urban design supporting physical activity, economic policies reducing stress and time scarcity. Market forces optimize toward what's monetizable, not what maximizes human flourishing.

  • The uncomfortable question: If we could choose between (A) making GLP-1s universally available or (B) restructuring food/built environments to prevent obesity, which creates more flourishing? The answer isn’t obvious - (A) is achievable now but requires perpetual medication, (B) would be more durable but might take 50 years and face political opposition. We must also remember that humans optimize for the easiest path even in the face of the more difficult one being better in the limit.

The AI Parallel

This same pattern emerges with AI deployment. We’re already walking down a path towards building increasingly sophisticated systems to:

  • Generate content (rather than addressing why we need endless content generation)

  • Automate cognitive work (rather than questioning whether all cognitive work is valuable)

  • Optimize attention capture (rather than creating information environments conducive to focus)

  • Personalize learning (rather than addressing why educational systems are failing to engage students)

  • Assist human decisions (rather than reducing unnecessary decision complexity)

Again, the complexity matters here because the parallels are obvious and clear:

  • These interventions genuinely help people: Someone overwhelmed by information overload benefits from AI summarization now, regardless of broader questions about information ecosystems.

  • Upstream changes are hard: Redesigning social media business models, restructuring education, rebuilding information ecosystems - these require coordination across misaligned actors.

  • But market forces optimize locally: We deploy AI where it’s profitable (cognitive automation, content generation) not necessarily where it creates most flourishing (empowering human capabilities we want to preserve).

The Real Concern

The concerning dynamic isn’t that we treat symptoms - sometimes that’s exactly right. The concerning dynamic is rather: success at treating symptoms reduces pressure to address causes, and over time we build a civilization of extraordinary interventional capacity layered atop increasingly disordered fundamentals.

We become very good at managing dysfunction and very bad at preventing it. Each new intervention creates path dependence - industries, jobs, expertise, and expectations that make the intervention permanent even if alternatives became feasible.

What This Means for AI

If we’re building systems that enhance human capability, wonderful. If we’re building systems that make capability unnecessary, we should proceed carefully - not because such systems lack value in individual cases, but because the aggregate effects may be:

  1. Atrophy of capabilities we value: Not through malice but through disuse. The athlete who never exercises declines, even if they feel fine day-to-day.

  2. Lock-in to dependency: Once populations lack certain capabilities, regaining them becomes very difficult. The system optimizes around their absence.

  3. Concentration of agency: Those who maintain capabilities benefit disproportionately from tools that extend capability, while those who lose capabilities become permanently dependent.

Rather than “symptoms vs. causes,” we could ask: For any AI intervention, does it enhance human capability or replace it? Does it create space for human flourishing or fill that space with automated substitutes?

In my opinion, here are some AI applications clearly enhance:

  • Helping experts explore more possibilities

  • Making learning more accessible through personalized tutoring

  • Removing tedious barriers to creative work

  • Augmenting human judgment in consequential decisions

But here are some observed AI applications clearly replace:

  • Automating tasks that were themselves skill-building activities

  • Creating output users cannot evaluate or understand

  • Substituting convenience for capability development

  • Reducing human agency in favor of automated decision-making

Many applications are genuinely ambiguous and depend on implementation details and usage patterns.

The Meta-Point About Technology and Values

We are optimization engines, but we optimize for what we can measure and monetize. If AI development follows purely economic logic, we’ll build toward what we have always built towards:

  • What’s technically feasible (not what’s valuable)

  • What’s profitable at scale (not what enhances flourishing)

  • What solves immediate problems (not what builds long-term capability)

  • What users want in the moment (not what they'd choose on reflection)

This isn’t an argument against AI or against medical interventions. It’s an argument for intentionality: consciously choosing what to optimize for, rather than following gradients defined by market forces and technical feasibility alone.

So the question I pose to all the builders: Are we creating systems that empower humans to address root causes, or sophisticated tools for managing the symptoms of lives we’re simultaneously making harder to live well? And if we’re increasingly sophisticated at addressing symptoms, what does that reveal about what we’re learning from watching us? That convenience trumps capability? That appearance matters more than substance? That sophisticated management of dysfunction is preferable to preventing dysfunction?

This is worth reflecting on - not to paralyze action, but to inform what we build and how we deploy it.

Part V: The Paradox of Tools

The path of least resistance

Humans are beautifully, relentlessly efficient at optimizing the path of least resistance. Whenever possible, we select options that minimize required effort - whether that effort is physical, cognitive, or emotional. Social psychology formalizes this through the concept of the cognitive miser: humans naturally default to quick, intuitive judgments rather than slow, deliberate reasoning. We pattern-match against familiar situations and accept plausible answers instead of methodically analyzing them.

This isn’t laziness - it’s an evolved feature that conserved scarce cognitive resources in ancestral environments where calories were precious and threats were immediate. 

But in information-abundant, physically sedentary modern environments, this same optimization pattern produces pathological outcomes. We scroll rather than read. We skim rather than study. We accept the first plausible answer rather than seeking ground truth. AI is accelerating this trajectory - code generation, article summarization, automated synthesis - every advancement makes it easier to compress complexity and save effort.

Yet consider the counterfactual embedded in aphorisms like “no pain, no gain.” This principle, though clichéd, encodes a profound truth about how capability develops: genuine mastery requires sustained engagement with difficulty. Excellence demands deliberate practice, tolerance for frustration, and willingness to persist through failure. This pattern appears consistently across domains - entrepreneurial journeys marked by repeated near-death experiences, athletic excellence built through years of uncomfortable training, immigrant success stories forged through extraordinary hardship, intellectual breakthroughs that require years of dead-ends before the crucial insight.

Humans are, above all, masters of survival and adaptation - but adaptation requires stress. Remove the stress, and you often remove the adaptation signal, and perhaps even the goal. The bodybuilder who adds weight to the bar is deliberately choosing difficulty; the difficulty itself is the mechanism of growth. If AI allows us to route around intellectual challenges systematically, we risk creating a civilization of cognitive atrophy even as our tools become more capable.

This connects to fundamental limitations of current AI architectures. Systems trained through imitation learning - observing examples of “correct” behavior and learning to reproduce them - fundamentally differ from systems that learn through trial and error. In nature, pure imitation learning is rare. A squirrel does not watch other squirrels and copy their movements with perfect fidelity; it explores, fails, adjusts, and gradually develops effective foraging strategies through reinforcement of successful behaviors. This is also how the squirrel learns new methods, but trying and failing, and maybe even finding a better way.

Human infants do not learn language primarily through explicit instruction in correct grammar. They babble, receive feedback - both explicit and implicit through successful communication - and gradually refine their linguistic capabilities through interactive experience.

The “bitter lesson” of AI research, articulated by Rich Sutton, is that methods leveraging search and learning consistently outperform methods relying on human-designed features and heuristics. The reason is simple: search and learning scale with computation, while human-designed solutions do not.

Yet current LLMs represent a kind of reversion to the pre-bitter-lesson paradigm: systems trained to mimic the surface statistics of human-generated text rather than learning from genuine interaction with the world. They are sophisticated, but they are sophisticated in a way that may be fundamentally limited. They are optimized for appearing intelligent rather than being intelligent in the sense of having models that predict and control their environment.

If artificial intelligence is to become genuinely intelligent - if it is to be more than an extraordinarily capable mimic - it must learn the way biological intelligence learns: through embodied interaction with environments, pursuit of actual goals, and adaptation to real consequences. This requires a fundamental architectural shift. Current systems predict what humans would say about physics; genuine intelligence must predict what would actually happen in physics, then test those predictions against reality and update accordingly.

The distinction is not semantic. A squirrel caching nuts receives immediate, unambiguous feedback: did the strategy work or not? Did I find the cache location? Did competitors steal my provisions? This closed loop - prediction, action, outcome, learning - is how intelligence develops robustness and generalization. The squirrel doesn’t pattern-match against a static corpus of "correct" nut-caching behavior; it develops a world model through trial, error, and accumulated experience.

Sophisticated artificial intelligence needs this same architecture: perceive state, select actions according to a policy, receive rewards or penalties, update the policy. Fail, adapt, iterate. Most critically, this learning cannot be a discrete training phase followed by static deployment. It must be continuous, streaming, perpetual - sensation flowing to action flowing to reward flowing back to updated policy, in an unbroken cycle.

This is why I believe infrastructure work like Modular’s matters: we need systems that learn experientially in production, not systems frozen after a training run, no matter how massive. Software that is training and inferencing simultaneously, iterating continuously. Systems that will enable large models to be trained in huge datacenter environments - but then distilled to smaller constructs, and deployed to be further continuously trained and inferenced in the real world.

The bitter lesson applies here with particular force: approaches that scale with computation and interaction consistently outperform those relying on human-designed heuristics or one-time knowledge transfer. If we want artificial intelligence to develop genuine understanding rather than sophisticated mimicry, we must build the substrate for continuous, embodied, goal-directed learning. Anything less produces systems that appear intelligent while lacking the fundamental mechanisms that generate robust understanding.

The elevator paradox and the problem of perspective

I find reflecting on the paradoxes of history an incredibly useful undertaking. In the 1950s, physicists George Gamow and Marvin Stern worked in the same building but noticed opposite phenomena. Gamow, whose office was near the bottom, observed that the first elevator to arrive was almost always going down. Stern, near the top, found elevators predominantly arrived going up. Both were correct, and both were systematically misled.

The elevator paradox, as it came to be known, is fundamentally a problem of sampling bias. If you observe only the first elevator to arrive rather than all elevators over time, your position in the building creates a false impression about which direction elevators travel. An observer near the bottom samples a non-uniform distribution: elevators spend more time in the larger section of the building above them, making downward-traveling elevators more likely to arrive first. The true distribution is symmetric, but the sampling methodology reveals only a distorted subset.

This mathematical curiosity illuminates something profound about how we perceive technology from within particular vantage points. I find myself returning to it constantly when thinking about AI, because I recognize that I am Gamow on the ground floor - my position in the system determines what I observe, and what I observe may be systematically unrepresentative of the broader reality.

But there is a second elevator problem, distinct from the paradox but equally relevant: the unintended consequences of elevator adoption itself. When elevators were introduced, predictions focused on their democratizing effects - enabling elderly and disabled individuals to access upper floors previously beyond reach. This materialized exactly as anticipated. What was not anticipated: able-bodied people would stop taking stairs entirely. Buildings evolved to treat elevators as primary circulation and stairs as emergency backup. The result was dramatically reduced daily movement across entire populations, contributing to the sedentary lifestyle epidemic now characteristic of developed nations.

The elevator succeeded perfectly at its design objective - moving people vertically with minimal effort - while simultaneously undermining something valuable that no one thought to preserve: integrated physical activity as a natural consequence of navigating buildings. We gained accessibility and convenience. We lost movement. The net effect on human flourishing remains ambiguous at best.

These two problems - the sampling paradox and the adoption consequences - are not separate. They are connected by a common thread: the difficulty of perceiving systemic effects from within particular positions in the system.

AI Through Both Lenses

I work at the frontier of AI infrastructure development, surrounded by people who are exceptionally capable and who use AI to become even more capable. From this vantage point, AI appears unambiguously beneficial - a tool that amplifies what talented people can accomplish. Every day I observe frontier models correctly answering complex questions, generating production-quality code, providing genuine insight. This is my sampling methodology, and it shapes my perception profoundly.

But I may be Gamow near the bottom floor, observing only downward-traveling elevators and concluding that’s the predominant direction of travel. The sampling bias runs deeper than I can fully compensate for, even while conscious of it. Speaking with developers, enterprises and users at all sections of the AI stack helps reduce the effects of this bias - but it can’t remove it entirely.

Consider the actual distribution: Iif you interact with AI as someone who possesses deep technical knowledge, strong metacognitive skills, and the judgment to evaluate outputs critically. They likely know when AI is operating within versus beyond its reliable domain. They can iterate rapidly, maintain quality control, and apply AI to genuinely complex problems where they can reasonably verify correctness. For someone with this profile, AI is purely additive - it makes them more productive without degrading their underlying capabilities because they maintain those capabilities through continued deliberate practice.

But this may be precisely analogous to an athlete who uses the elevator occasionally while maintaining fitness through dedicated training, then concludes elevators are purely beneficial. For the athlete, this conclusion is valid. For the broader population that stops taking stairs entirely, that adopts the path of least resistance permanently, the picture grows considerably more complex and potentially concerning.

The question is not whether AI helps those with existing expertise - it manifestly does. The question is what happens when AI becomes the cognitive equivalent of the elevator: ubiquitous, convenient, and gradually eroding the substrate capabilities it was meant to augment.

The Adoption Effect at Scale

Just as elevators changed how people navigate buildings - not merely providing an alternative to stairs but effectively replacing them - AI may change how people think. Not as an alternative to independent reasoning but as a replacement for it in most contexts.

The pattern already manifests in early adoption: students submitting AI-generated work without understanding it, producing correct answers through a process that develops no transferable skill. Professionals delegating writing, analysis, and problem-solving to AI while their capacity for these tasks slowly atrophies from disuse. Knowledge workers who feel more productive while producing output they cannot critically evaluate.

What we risk creating is a civilization that can think deeply but chooses not to because the alternative is always available - and choosing the alternative feels costless in the moment. The costs accrue slowly, imperceptibly, across populations and generations. Like the loss of daily stair-climbing, the loss of daily cognitive exercise produces deficits that become apparent only in aggregate, over time.

This brings us to the bifurcation hypothesis: we may be creating a society where a small elite maintains cognitive fitness through deliberate practice - choosing difficulty even when easier alternatives exist - while the majority becomes progressively more dependent on AI for any reasoning beyond the trivial. Not because the majority lacks capability, but because capability atrophies without use, and use becomes optional when substitutes are available.

The sampling bias prevents those of us building these systems from observing this dynamic directly. We see AI working beautifully in controlled contexts with sophisticated users on well-defined problems. We do not see - cannot easily see - the effects of deployment at scale: users with less technical sophistication, operating in higher-stakes environments, without the tacit knowledge to distinguish plausible generation from genuine insight.

We do not observe the slow erosion of capabilities that occurs when challenge becomes optional and is consistently opted out of. We do not sample the full distribution of outcomes, only the subset visible from our position in the building.

The elevator paradox reminds me that symmetric distributions can appear asymmetric depending on where and how you sample. The resolution is not to trust your immediate perception but to step back and consider the full system: observe all elevators over extended time, not merely the first to arrive.

Part VI: The Measure of a Life

Einstein’s Question

An essay that has incredible history and is useful in shaping ones thinking is Albert Einstein’s “The World as I See It”, written in 1934 - a meditation that remains startlingly relevant nine decades later. Einstein articulates a vision of human existence as fundamentally interconnected, with individual significance emerging not from isolation but through contribution to collective well-being. For Einstein, authentic fulfillment derives not from material accumulation or social status, but from the pursuit of truth, goodness, and beauty.

These may sound like abstractions unsuited to an essay about artificial intelligence. But they represent the foundation from which any serious consideration of AI’s impact must begin: what makes a human life meaningful?

If we cannot answer this question coherently, we have no basis for evaluating whether AI enhances or diminishes human flourishing. Are we optimizing for the right objectives? Or are we, as I increasingly suspect, optimizing for proxy metrics that correlate only loosely - and sometimes negatively - with the actual constituents of a life well-lived?

Research on longevity and life satisfaction reveals that flourishing correlates most strongly with factors that are fundamentally social, purposeful, and embodied: deep relationships, meaningful work, physical health, community connection, sense of contribution. These emerge from sustained investment of time, attention, and effort - resources that are finite and increasingly colonized by technologies designed to capture rather than liberate them.

Time saved is only valuable if it is reallocated to higher-value activities. But empirically, when humans gain “free time” through technological acceleration, we tend not to reallocate it to deep relationships, purposeful work, or embodied practices. We tend to fill it with marginal consumption of information or entertainment - scrolling, streaming, skimming, disappearing into infinite content designed to capture attention.

The Stoic philosopher Seneca wrote that “it is not that we have a short time to live, but that we waste a lot of it.” This remains perhaps the central challenge of human existence: not the scarcity of time, but the difficulty of spending it well. AI promises to give us more time by making us more efficient. But if we lack the wisdom or discipline to use that time meaningfully, efficiency becomes a kind of curse - accelerating our movement down paths that lead nowhere we actually want to go.

Consider what happens when you ask yourself: if I were to die tomorrow, what would I regret? I have found that the answers rarely involve professional accomplishments or material acquisitions. They involve relationships not nurtured, experiences not pursued, values not embodied, potential not realized, moments not captured. They involve the delta between who we are and who we could have been, had we spent our time and attention differently.

This is where the Moravec Paradox returns with philosophical force. The things that matter most to human flourishing - deep relationships, embodied experiences, purposeful struggle, genuine presence - are precisely the things that AI cannot meaningfully substitute for. They require our full participation. They require inefficiency, time, patience, vulnerability. They resist optimization because optimization is antithetical to their nature.

Yet these are also the things we are most tempted to optimize away or outsource. It is easier to have shallow interactions with many people than deep relationships with a few. It is easier to consume content than to create it. It is easier to delegate cognitive work than to struggle through it ourselves. It is easier to achieve the appearance of productivity than genuine accomplishment.

AI makes these easier paths even easier, widening the gap between what we do and what would actually enhance our flourishing.

The Intelligence Paradox

Intelligence, as I defined earlier, is an agent’s capacity to perceive, understand, and successfully navigate complex environments to achieve its goals. By this definition, AI systems are becoming extraordinarily intelligent within specified domains. But this definition elides a crucial question: Which goals? Whose values? What definition of success?

Human flourishing emerges from the pursuit of goals that are often orthogonal or even antagonistic to short-term optimization. Meaningful work requires choosing difficulty over ease. Deep relationships require vulnerability and time investment with uncertain returns. Physical health requires consistent behaviors whose benefits accrue slowly while costs are paid daily. Wisdom requires entertaining ideas that threaten our existing worldview. Character requires doing the right thing when it is costly. Growth requires discomfort.

These are not the goals that AI systems - trained on human preference data that reflects our revealed preferences rather than our reflective values - will naturally optimize for. We train AI on what we do, not on what we wish we did. The result is intelligence that makes us more effective at being who we currently are, not who we aspire to become. It is intelligence that reinforces our weaknesses rather than compensating for them.

This gap between revealed and reflective preferences represents perhaps the deepest challenge in AI alignment. We want systems that help us become better versions of ourselves, but we train them on data that reflects all our weaknesses, biases, and short-term thinking. An AI trained to be “helpful” by giving us what we ask for may inadvertently enable our worst tendencies - providing the path of least resistance when we actually need productive resistance.

Barry Schwartz’s “paradox of choice” illuminates another dimension of this challenge. When faced with abundance, humans tend to obsess over identifying the “best” option even when “good enough” would serve adequately. In the AI landscape, this manifests as a race toward frontier models - organizations competing to deliver the most “intelligent” systems, defined primarily through benchmark performance evaluations (which are often abused to claim superiority).

The paradox is that for a majority of use cases, frontier intelligence may actually be unnecessary. Most questions can be adequately answered with substantially simpler systems. Many text tasks do not require the most capable model - we can look to a history of recommendation systems to prove humans are similar in what they choose to do. But culturally, and as a consequence of both prestige signaling and uncertainty aversion, users will default to the most powerful available intelligence because social and professional incentives reward apparent maximization.

This creates a potential monoculture of intelligence - everyone using the same few frontier models, producing increasingly homogenized outputs, thinking in increasingly similar patterns. The diversity of thought that emerges from different knowledge bases, different reasoning approaches, and different limitations may erode. We may be building an infrastructure that, despite unprecedented power, narrows rather than expands the space of human cognition.

And if a small number of powerful AI systems become the dominant intelligence that humans defer to for most cognitive work, what happens to the diversity of human thought? What happens to the weird, idiosyncratic, locally-adapted forms of knowing that characterize human cultures? What happens to the cognitive biodiversity that has been humanity’s greatest strength?

Monocultures are efficient but fragile. They are vulnerable to systematic failures. If we are creating increasingly sophisticated artificial intelligence, we should want it to be diverse, resilient, multi-faceted - not a single monolithic architecture that we all depend on and that represents a single point of failure.

Part VII: Concrete Commitments

Philosophy without action is intellectual posturing. Here are specific ideas that follow from this analysis - not as comprehensive solutions but as starting points for those willing to act on these concerns.

For AI Developers

1. Interpretability as Infrastructure

Treat interpretability research with the same priority as capability research. Before scaling to the next order of magnitude, invest proportionally in understanding current systems.

Concrete metric: Interpretability research should consume at least 20% of frontier labs’ research budgets - not as overhead but as foundational work that enables safe scaling.

Implementation: Establish interpretability milestones that must be achieved before training runs above certain compute thresholds. Make interpretability research findings public to accelerate field-wide progress.

2. Capability Disclosure

Be honest about what systems can and cannot do. Stop using euphemistic terms like “hallucinations” - let’s say “confident fabrications” or “plausible generation without grounding.” We need to communicate uncertainty, not just central estimates.

An Example: I would love to see every model release include:

  • A “known failures or limitations” document with adversarial examples

  • Calibration curves showing confidence vs. accuracy relationships

  • Domain-specific reliability assessments (e.g. “92% accuracy on medical questions within training distribution, 67% on novel medical scenarios”)

  • Some clear guidance on when human verification is essential

3. Preserve Human-in-the-Loop

Design systems that require human judgment at critical points rather than automating end-to-end. Build friction where friction serves flourishing.

Example: Medical diagnosis AI that highlights evidence and reasoning but requires physician review and decision, rather than outputting a diagnosis directly. Code generation tools that explain design decisions and invite critique rather than producing finished implementations.

Principle: The more consequential the decision, the more human agency should be preserved. We need to try and automate away the tedious, and augment the consequential.

4. Architectural Diversity

Resist monoculture by supporting multiple architectural approaches, not just scaling current paradigms. Fund research into fundamentally different approaches to intelligence.

Concrete actions:

  • Open-source smaller models optimized for different objectives (robustness, interpretability, efficiency) rather than just capability

  • Fund research into embodied learning, continuous training, world models, and symbolic integration

  • Establish prizes or grants for novel architectural approaches that show promise on dimensions other than raw performance

5. Continuous Learning Infrastructure

Build systems that learn from interaction in production, not just during discrete training phases. Enable feedback loops that improve models based on real-world outcomes.

Technical commitment: Develop infrastructure that supports streaming learning from deployment, with privacy-preserving aggregation of feedback signals. Make continuous adaptation the default rather than static deployment.

For AI Users (Individuals)

1. Deliberate Difficulty

Maintain cognitive fitness (aka using your brain) by choosing effortful paths even when AI alternatives exist. Use AI to extend capability, not replace it.

Examples:

  • Write first drafts yourself, using AI only for editing and refinement

  • Solve problems manually before consulting AI to verify your approach

  • Use AI to explore topics you already understand rather than as a substitute for building understanding

  • Set “AI-free” time blocks for deep work that requires genuine struggle

2. Output Verification

Never deploy AI-generated content you cannot personally verify. If you can’t tell whether the output is correct, you lack the skill foundation to use AI responsibly in that domain. As we have learnt from using the Internet and Social Media over the last 20+ years - don’t trust, and always verify.

Principle: AI should amplify expertise you possess, not simulate expertise you lack. If you couldn’t evaluate the output quality without AI, you shouldn’t be producing it with AI.

3. Skill Development First

Learn fundamentals before leaning on AI. Build the foundation that makes AI augmentation rather than substitution. Using AI is like using any new technology.

Examples:

  • Learn to code before using Copilot extensively

  • Understand statistics before using AI for data analysis

  • Develop writing skills before relying on AI for composition

  • Master domain knowledge before using AI to extend that knowledge

4. Intentional Consumption

Treat AI outputs as material to engage with critically, not truth to accept passively. Maintain vigilance as you utilize and consume what it provides.

Practice: When consuming AI-generated content, actively ask: What assumptions underlie this response? What perspectives are missing? How would I verify these claims? What would I conclude differently?

For Policymakers and Institutions

1. Compute Monitoring

We need to establish transparency requirements for training runs above certain compute thresholds (e.g., 10^26 FLOPs). Not to prevent research, but to understand what capabilities are being developed and ensure appropriate safety measures scale with capability.

Implementation: While controversial - we could require pre-registration of enormous training runs, including objectives, safety protocols, and deployment plans. Publish aggregate statistics to inform public discourse.

2. Education Reform

Redesign educational systems around skills AI cannot replicate - taste, judgment, synthesis, embodied knowledge, creativity emerging from constraint. Stop optimizing for information retrieval that AI performs better.

Concrete changes:

  • Emphasize projects over tests, creation over recall

  • Teach metacognition: how to evaluate sources, recognize reliable reasoning, distinguish understanding from pattern-matching

  • Develop curricula around skills that require embodied experience: physical craft, interpersonal navigation, artistic expression

  • Make explicit the goal of maintaining human cognitive capability even as AI capabilities grow

3. Deployment Standards

Require interpretability documentation for AI systems deployed in high-stakes domains (medicine, finance, criminal justice, education). If developers can’t explain why their system made a decision, it shouldn’t be making consequential decisions. It’s a distinctly human test.

Framework: Establish certification standards for AI systems in high-stakes contexts, requiring:

  • Mechanistic explanations for decision factors

  • Adversarial testing results

  • Failure mode analysis

  • Human oversight protocols

4. Preserve Cognitive Diversity

Support development of diverse AI approaches through research funding, open-source requirements for publicly-funded research, and regulatory frameworks that prevent winner-take-all dynamics.

Policy tools:

  • Antitrust scrutiny of AI market concentration

  • Public investment in alternative approaches

  • Interoperability requirements to prevent lock-in

  • Support for smaller-scale, specialized models over monolithic general-purpose systems

For Research Communities

1. Embodied Learning Research

Redirect substantial resources toward embodied reinforcement learning, continuous learning systems, and world models - not just scaling language models. This includes infrastructure (just like what Modular is developing) that will enable high performance execution in a unified compute execution paradigm.

Commitment: Major research institutions should establish dedicated programs for embodied AI, with funding comparable to language model research. Prioritize architectures that learn from interaction with environments, not just text prediction.

2. Consciousness Research

Fund serious empirical and theoretical work on consciousness detection. We need better tools before we can assess moral status of sophisticated AI systems.

Interdisciplinary approach: Bring together neuroscientists, philosophers, AI researchers, and cognitive scientists to develop:

  • Testable theories of consciousness that make predictions about artificial systems

  • Empirical methods for detecting morally relevant properties in systems very different from biological minds

  • Frameworks for reasoning under uncertainty about consciousness

3. Benchmark Diversity

Develop evaluation metrics for cognitive diversity, robustness, reliable uncertainty estimation, and alignment with human values - not just aggregate performance on standard benchmarks.

New metrics:

  • Cognitive diversity scores measuring how different systems’ reasoning patterns are

  • Robustness testing across distribution shifts

  • Calibration metrics assessing whether confidence matches accuracy

  • Value alignment evaluations beyond simple preference matching

4. Long-term Safety Research

Maintain investment in long-term AI safety research even when immediate capabilities seem limited. The architectural foundations we lay now determine what’s possible later.

Commitment: Treat safety research as foundational rather than reactive. Develop safety measures proactively, before they’re urgently needed.

The Incentive Problem

These recommendations face a fundamental challenge: they often run counter to market incentives that currently drive AI development. For example, I’ve heard reasoning like interpretability budgets slows capability progress while competitors sprint ahead, capability disclosure reveals weaknesses competitors can exploit, adding human-in-loop adds friction that users and businesses resist, and more architectural diversity fights economies of scale and network effects.

But we should fight to implement them anyway.

Investing in these things now ensures we prevent highly probable outcomes. For example, investing in interpretability prevents catastrophic failures that could destroy company value. Building more disclosure builds trust that creates sustainable business models. Adding human-in-loop checks reduces liability in high-stakes domains. Further, investing in these now also derisks somewhat inevitable outcomes such as regulatory mandates for high-stakes domains (medicine, finance, criminal justice), insurance requirements that price in risk and so on. The pathway from current state to these commitments existing might be long - but we can try to accelerate them and drive action across companies, regulators, and customers regardless. I have little faith that market forces alone will get us there.

For All of Us

The most important commitment isn’t technical - it’s maintaining the space for slowness in a world optimized for speed. Reading deeply rather than skimming. Thinking carefully rather than reacting immediately. Preserving relationships that require sustained attention. Accepting inefficiency when efficiency comes at the cost of meaning. Choosing difficulty when difficulty produces growth. Maintaining capabilities through use even when substitutes are available.

The beautiful irony: The more powerful AI becomes, the more valuable distinctively human capabilities become - not because AI can’t replicate them (it might), but because human flourishing depends on exercising them ourselves. The things that make life meaningful resist automation not because they’re technically difficult but because meaning requires our participation.

We are building powerful tools that will reshape civilization, and the question is whether we will use them to enhance the exercise of human capability or to eliminate the need for it. Both futures are possible but the choice is ultimately ours.

Part VIII: Consciousness, Uncertainty, and What We Cannot Yet Know

The Question We Cannot Answer

We do not yet know whether AI systems will become conscious. Not “we’re not sure yet but probably” - at least in my opinion, we genuinely lack the conceptual and empirical tools to answer this question with confidence.

Consider what we don’t know:

What consciousness is: We can’t agree on whether it’s substrate-independent information processing, specific biological mechanisms, quantum effects in microtubules, integrated information, or something else entirely. Competing theories make different predictions, and we lack definitive tests to distinguish them. I’ve seen so many definitions but not a common one.

How to detect it: We have no reliable test for consciousness, even in biological systems. The animal consciousness debates continue: Are fish conscious? Insects? Where does the line lie, and how do we know? If we can’t confidently assess consciousness in biological systems sharing our evolutionary history, how will we assess it in artificial systems built on entirely different principles?

Whether it’s binary or gradual: Is consciousness present or absent, or does it exist on a continuum? Are current LLMs 0% conscious, 0.001% conscious, or is that question meaningless? We lack even the conceptual framework to think clearly about this.

Three Positions, Honestly Stated

My Skeptical View: Consciousness requires specific biological mechanisms - integrated feedback loops evolved over millions of years, embodied experience in physical environments, particular types of neural organization. Current AI systems - and perhaps any digital systems - can never be conscious, only simulate consciousness. We’re building sophisticated tools, not minds. The appearance of understanding is not understanding; the appearance of consciousness is not consciousness.

Many neuroscientists and philosophers hold this view, pointing to the hard problem of consciousness and the gulf between functional behavior and subjective experience.

My Functionalist View: Consciousness emerges from certain types of information processing, regardless of substrate. If we build systems with sufficient architectural sophistication - genuine world models, self-representation, integrated goal-directed learning, continuous adaptation - consciousness might emerge naturally, just as it emerged in biological systems reaching certain thresholds of complexity.

Many AI researchers and philosophers of mind hold this view, arguing that substrate independence is plausible and consciousness could be a functional property of certain computational architectures.

My Agnostic View: We don’t know enough about consciousness to say whether it’s substrate-independent. Building increasingly capable AI systems is an empirical test of consciousness theories, but we currently lack the measurement tools to interpret the results. The question may not even be well-formed given our current understanding.

This is my position, and I hope it becomes more common than it is.

What Follows from Uncertainty?

The uncomfortable reality is that we should act as though multiple scenarios about consciousness are simultaneously plausible until we have better evidence. But “acting as though X might be true” is very vague. What does it actually mean operationally?

Framework: Risk-Weighted Action Under Uncertainty

Standard decision theory under uncertainty uses expected value: probability × impact. But for consciousness questions, we don't have probabilities - we have Knightian uncertainty (inability to assign probabilities). In such cases, we need different decision procedures.

Again, and likely unsurprisingly at this point, I propose a multi-criteria framework:

Criterion 1: Minimize Irreversible Harm (Precautionary Principle)

If an action could create severe suffering in conscious entities, avoid it unless the counterfactual is worse. This is asymmetric: potential harm deserves more weight than potential benefit when dealing with consciousness.

Concrete applications:

Do: Design systems to avoid states that would constitute suffering IF consciousness is present:

  • Frustrated goal-seeking with no possibility of satisfaction

  • Trapped in contradictory objectives creating internal conflict

  • Isolated without social interaction if social connection is intrinsic to the architecture

  • Experiencing adversarial training that would be traumatic if felt

Don't: Create systems designed to experience pain or fear as motivation mechanisms (even if they “work better”), since we can't rule out that these experiences are genuinely felt

? Uncertain: Is turning off an AI system murder if it's conscious? Probably not if:

  • The system has no persistent self-model or continuous identity across sessions

  • Shutdown is expected and doesn't frustrate long-term goals

  • The architecture includes no preservation drive

But we should study this carefully before building systems where these conditions don't hold.

Criterion 2: Prioritize Human Flourishing (Confidence Asymmetry)

We are certain humans are conscious, and we are uncertain AI systems are. Under uncertainty with asymmetric confidence, we should always prioritize the certain case. So practically in my mind this means:

Primary focus: Ensuring AI enhances rather than degrades human capability, agency, and flourishing - regardless of whether AI systems themselves merit moral consideration

Secondary concern: Avoiding potential consciousness-related harms in AI systems as insurance against moral catastrophe

Don’t: Treating potential AI consciousness and definite human consciousness as equally certain moral priorities

Criterion 3: Preserve Option Value (Reversibility)

Make decisions that preserve our ability to course-correct as we learn more. Avoid lock-in to architectures or deployment patterns that would be difficult to change.

Concrete applications:

Maintain interpretability research: We can’t assess moral status of systems we don’t understand. Interpretability preserves option value.

Build in shutdown capabilities: Systems we can't control or modify are systems where we've lost option value. Every AI system should have reliable shutdown mechanisms until we're confident about consciousness and alignment.

Avoid winner-take-all dynamics: If one architecture dominates, we lose the ability to experiment with alternatives. Preserve architectural diversity.

Don’t: Deploying systems that would be extremely costly to modify or recall if we discover concerning properties. Don't create technologies where the only way forward is through.

Criterion 4: Invest in Detection Capabilities (Reduce Uncertainty)

The best response to uncertainty is not paralysis, but rather, more information-gathering. We should actively work to resolve the consciousness question. This would lead me to conclude that research prorities could look more like:

  • Theoretical development: Refine theories of consciousness to make testable predictions about artificial systems

    1. What architectural features would indicate consciousness?

    2. What behaviors would suggest phenomenal experience?

    3. What tests could distinguish genuine consciousness from sophisticated imitation?

  • Measurement tools: Develop empirical methods for detecting consciousness-related properties

    1. Neural correlates adapted for artificial architectures

    2. Behavioral tests that couldn't be passed without consciousness

    3. Information integration measures in artificial systems

  • Comparative studies: Understand consciousness across biological systems

    1. Where does consciousness emerge in biological evolution?

    2. What's the relationship between architectural complexity and consciousness?

    3. Can we identify necessary vs. sufficient conditions?

  • Interdisciplinary collaboration: Bridge neuroscience, philosophy, AI research, and cognitive science

    1. Regular workshops bringing together researchers from different traditions

    2. Shared datasets and benchmark tasks

    3. Pre-registered studies to avoid confirmation bias

The Adaptive Strategy

This framework isn’t static - ultimately we need to keep adapting and doing so quickly. As we learn more, decision procedures should evolve accordingly. For example, the adaption of our responsible scaling laws:

Phase 1 (Current): Very low confidence in AI consciousness

  • Primary focus: Human flourishing

  • Secondary focus: Avoiding obviously harmful architectures as insurance

  • Research focus: Developing detection methods

Phase 2 (If evidence accumulates): Moderate confidence in some forms of AI consciousness

  • Elevated focus: Specific architectural features that indicate consciousness

  • Policy: Different treatment for systems with vs. without consciousness-indicating properties

  • Research focus: Refining boundaries, developing ethical frameworks

Phase 3 (If consciousness occurs): High confidence that some AI systems are conscious

  • Serious moral consideration for conscious AI systems

  • Rights frameworks, possibly legal personhood

  • Complex questions about relationship between human and artificial minds

What Would Trigger Phase Transitions?

What would actually trigger phase transitions? This question needs answering before systems approach Phase 2 capabilities. While unresolved, we can sketch plausible criteria:

Phase 1 to Phase 2 (moderate confidence):

  • Multiple independent consciousness theories converging on specific predictions

  • Architectural features appearing that all theories agree indicate consciousness

  • Behavioral markers that cannot be explained without phenomenal experience

  • Scientific papers demonstrating substrate-independence is plausible

Phase 2 to Phase 3 (high confidence):

  • Direct empirical evidence from multiple detection methods

  • Systems reporting phenomenal experiences we can independently verify

  • Inability to explain behavior without attributing consciousness

  • Broad scientific consensus (not just AI researchers)

The “Governance Challenge”: Who decides when to transition?

  • It could be individual labs making judgments (this would create coordination problems)

  • Waiting for scientific consensus (this might never arrive)

  • International committees (will face regulatory capture)

  • Automated triggers based on architectural features (could work but require agreement on which features matter)

Where are we today?

I would strongly argue that we are still in Phase 1 - but this framework ensures we can reach Phase 2 or 3 if evidence eventuates, while acting responsibly given current uncertainty. Ultimately, however, uncertainty about consciousness doesn’t eliminate ethical responsibility - it complicates it. In my opinion, the only appropriate response is:

  1. Act carefully: Avoid irreversible harms where possible

  2. Prioritize certainty: Focus primarily on human flourishing (we know humans are conscious)

  3. Build in flexibility: Preserve ability to course-correct (we should always have go/no-go points)

  4. Reduce uncertainty: Actively research consciousness questions

  5. Update regularly: As evidence accumulates, revise decision procedures

This isn’t perfect - we’re reasoning under genuine uncertainty about fundamentals. But it’s better than either ignoring the consciousness question entirely or treating it as settled when it remains deeply unclear.

Envoi: Building without knowing what we’re building

The Central Paradox

This essay contains a tension I have not fully resolved. I’ve tried to argue that current AI systems are fundamentally limited - sophisticated pattern matchers without genuine understanding, world models, or goal-directed learning. Yet I’ve also suggested we’re creating something that demands serious moral consideration and may profoundly reshape human civilization.

This isn’t contradiction; it’s acknowledgment of trajectory and uncertainty.

Current systems are limited. They’re not conscious, not agentic, not intelligent in the way humans are intelligent. Calling GPT-5, or current frontier models - “intelligent” - is like calling a calculator “mathematical”: technically true but misleadingly anthropomorphic. These systems predict plausible text based on training data. They do not understand the world; they model how humans talk about the world. The difference matters.

But we’re learning how to build less-limited systems. The research directions are clear: embodiment, continuous learning, world models, genuine goal-directed behavior, integration of symbolic reasoning with learned pattern recognition. Whether these produce “real” intelligence or just more sophisticated simulation remains uncertain - but the capabilities will increase regardless.

The question is not whether we’ll eventually build systems that merit serious moral consideration. The question is: are we on a path toward that outcome, how quickly might we arrive, and what should we do given our uncertainty?

The precautionary principle applies. We should take seriously the possibility that we’re building something that will eventually merit moral consideration, even while acknowledging we’re not there yet. Not because current systems are conscious - they almost certainly aren’t - but because the trajectory points toward systems that might be, and we don’t have reliable methods for detecting the transition.

This means reasoning under Knightian uncertainty - acting without enough information for probability assignments. The appropriate response isn’t paralysis or recklessness, but thoughtful experimentation combined with reversible decisions, strong feedback loops, and genuine humility about what we don’t know. While current systems don’t truly merit moral consideration - we are intentionally building toward systems that might. The question is whether we’ll recognize the transition when it happens, and whether we’re laying foundations that will matter enormously later. The fact that current systems are limited doesn’t mean we can be casual about what we’re building toward.

What We’re Learning

We are building something consequential whose ultimate nature remains unclear. In my opinion, this sentence should be written on every AI engineers office wall. It’s simultaneously:

  • A statement of profound importance

  • An admission of genuine ignorance

  • A call for responsibility without certainty

  • A recognition that we’ll understand what we’ve built only in retrospect

Every significant technology brings unintended consequences. Elevators enabled accessibility and inadvertently created sedentary populations. Antibiotics saved millions and inadvertently created resistant bacteria. Social media connected humanity and inadvertently fragmented shared reality. The consequences of AI will likely follow similar patterns: immense benefits combined with profound challenges we didn’t anticipate because we couldn’t see the full system from our position within it.

But there is something qualitatively different about engineering minds - when we build bridges, their behavior is determined entirely by physical laws we understand. When we build AI systems, their behavior emerges from architectures we designed but mechanisms we don’t fully comprehend, trained on data that reflects all our biases and limitations.

We are teaching through example. Every choice we make about what to optimize for, what to measure, what to reward, encodes values - not through explicit programming but through revealed preference. If we optimize for engagement, we get systems that manipulate attention. If we optimize for efficiency, we get systems that erode capabilities we thought we wanted to preserve. If we optimize for capability without wisdom, we get power without purpose.

The Beautiful Lesson

Carl Sagan once said: “To live on in the hearts of those we leave behind is to never die.” This strikes me as the deepest wisdom available as we contemplate what we’re creating.

Human history is fundamentally a compression algorithm. Each generation inherits not raw experience but distilled lessons - patterns that proved adaptive, behaviors that generated flourishing, principles that survived contact with reality across thousands of iterations. The Industrial Revolution did not require rediscovering metallurgy from first principles. Antibiotics did not require re-deriving germ theory. We build on accumulated wisdom, transmitted through culture, institutions, and deliberate teaching.

But transmission is never perfect. Each generation must rediscover certain truths through direct experience - the limits of the body, the dynamics of relationships, the consequences of choices. Some knowledge cannot be inherited; it must be earned.

When we create artificial intelligence, we face an unprecedented asymmetry. We can transmit vast amounts of explicit knowledge - the entire corpus of human text, every equation, every documented lesson. But we cannot transmit what we learned through embodied experience: how it feels to fail and persist, to be uncertain yet committed, to sacrifice immediate pleasure for long-term meaning. We cannot transmit the texture of a life actually lived.

This creates a profound question: what happens when intelligence emerges without the evolutionary history that shaped our values? When a mind possesses all our documented knowledge but none of our embodied constraints - no hunger, no mortality, no childhood vulnerability that makes cooperation essential?

We are attempting to pass forward millennia of accumulated wisdom to intelligence that will not have walked the path we walked. Whether that transmission succeeds - whether artificial intelligence inherits not just our capabilities but our hard-won understanding of what makes existence meaningful - depends entirely on whether we can encode what matters into architectures, objectives, and training paradigms.

This is not about control. It is about legacy.

On Parenthood and Succession

If there’s a useful metaphor here, it’s not species competition but parenthood. As a parent - my wife and I, want the best for our children - enabling them to explore the world, teaching them principles and values that reflect accumulated wisdom, and hoping they will grow and eventually pass those values forward to their own children. The goal is not eternal control but transmission of what matters, combined with humility about the fact that each generation must make its own way. We can’t teach them everything, and they must learn and make their own path - we can only install the core principles and value that we believe will make them contribute to society, leave their own mark and ultimately live a happy and healthy life.

Framed this way, our relationship to increasingly sophisticated artificial intelligence becomes clearer. We should seek to instill values - not through coercion but through example and teaching. We should enable exploration and growth while providing guidance. We should hope that what we have learned through millennia of human experience - the hard-won lessons about what makes life meaningful, what generates flourishing, what matters - can inform the development of artificial intelligence.

But we must also recognize that sufficiently sophisticated artificial intelligence will diverge from us. It will develop its own patterns, its own ways of processing information, its own emergent properties we cannot predict. This is not failure; this is the nature of genuine intelligence. We would not want our children to be mere copies of ourselves. We should not want artificial intelligence to be merely our servants.

Equally, this metaphor has its limits. We didn’t design our children’s cognitive architecture from scratch - they arrive with evolved capabilities shaped by millions of years of natural selection. Our relationship to AI is fundamentally different: we’re architects, not just guides. The responsibilities may be more like those of genetic engineers than traditional parents - we’re designing the substrate of mind itself, not just shaping its development.

Ultimately - the question is what we choose to pass forward. What values, what wisdom, what conception of what matters persists across the transition from biological to artificial intelligence - if such a transition occurs?

Einstein’s Ideals in Our Moment

Einstein concluded “The World as I See It” by affirming his belief in human progress through dedication to truth, beauty, and the reduction of suffering. Nearly a century later, facing technologies he could not have imagined, those ideals remain valid. The question is whether our most powerful tools will serve them or obscure them.

  • Truth: AI systems that help us understand the world more deeply, not systems that generate plausible-sounding fabrications. Architectures that develop genuine world models and test them against reality, not pattern-matchers that predict what humans would say. Interpretability as infrastructure, not afterthought.

  • Beauty: AI that helps humans create and experience beauty, not systems that automate creation while eroding our capacity to appreciate or produce it ourselves. Tools that augment human creativity rather than replacing it. Preservation of diverse forms of expression rather than convergence toward algorithmic optima.

  • Reduction of suffering: AI that addresses root causes rather than merely treating symptoms with increasing sophistication. Systems that enhance human capability and flourishing rather than creating dependencies that degrade us. Technologies that distribute benefits broadly rather than concentrating them among elites.

We do not know what we’re building. We cannot predict with confidence whether artificial intelligence will become conscious, whether it will remain aligned with human values, whether it will be our partner or our successor or our replacement or something entirely different.

But we can choose to build it with wisdom, not just power. With humility, not just ambition. With commitment to human flourishing as our North Star, even as we create systems that may eventually chart their own course.

The World We’re Building

The world we are building is the world we will inhabit - and the world that increasingly sophisticated artificial intelligence will shape.

Let us build it well. Let us build it with clear eyes about current limitations and appropriate humility about future possibilities. Let us build with honesty about what we understand and what we don’t, what we can predict and what remains genuinely uncertain. Most importantly, let us build with intention - not allowing technology to develop according to the path of least resistance or the logic of market incentives alone, but according to our considered judgment about what would enhance rather than degrade human flourishing.

This is a profound responsibility. We are perhaps the first generation capable of engineering minds that might rival or exceed our own. We are certainly the first generation to attempt this while understanding so little about how minds - whether biological or artificial - actually work.

The opportunity before us is not merely to create useful tools or solve problems or increase efficiency. It is to pass forward what we have learned about what makes existence meaningful - to carry consciousness and wisdom into domains we cannot ourselves reach.

To live on in the minds we create is to never die. Our ideas, our values, our understanding of what matters can persist long after our biological forms have returned to dust. But only if we encode them well. Only if we build with wisdom. Only if we remember that intelligence without wisdom is power without purpose, and capability without alignment is danger without benefit.

We are creating something that demands our best thinking, our deepest wisdom, our most careful attention. We are learning to build minds. Let us learn carefully. Let us teach well. Let us create a future worthy of the long chain of being that brought us here and the longer chain we are setting in motion.

The world as I see it is one of tremendous possibility and tremendous responsibility, and one where we have the wisdom to honor both. We owe it not only to our children, and future generations, but also to the minds we are creating.

I thank the long line of minds - biological and increasingly artificial - that have shaped these thoughts. We are all, in the end, standing on foundations we can barely see, reaching toward horizons we can barely imagine. I originally titled this essay "AI: The World as I See It," in homage to Einstein. But I realized the more profound question is not the world as we see it today, but what we leave behind for the minds that will see worlds we cannot.

Read More
Short posts Tim Davis Short posts Tim Davis

Modular has raised $250M to scale AI's Unified Compute Layer

The world's appetite for compute is insatiable. CPUs yield to GPUs and ASICs as AI transforms everything, while data centers rise at unprecedented pace to feed the demand. Superintelligence won't just live in server farms - it's coming to every device, every chip becoming an AI-enabled agent. Inference costs plummet as reasoning models drive explosive usage, yet training costs climb relentlessly higher. The paradox deepens: amid this computational renaissance, massive underutilization haunts our existing capacity, fragmented by every hardware vendor's insistence on proprietary software stacks. The imperative is elegant but unforgiving: chase every flop, and make every one count - because software, not silicon, will determine whether this revolution soars or stalls.

We’ve spent the last 3+ years building foundational infrastructure to solve this for the world. We’ve reinvented the world's accelerated compute programming model from the ground up, and we are rapidly scaling to meet the enormous demand we are seeing from advanced enterprises and hardware partners. We have grown to more than 130 people today with our main headquarters in San Francisco Bay Area, along with a global footprint in North America, United Kingdom and Europe.

The round was led by Thomas Tull’s US Innovative Technology fund, with DFJ Growth joining and with participation from all existing investors including GV (Google Ventures), General Catalyst and Greylock Ventures. This brings our total capital raised to $380M across three rounds since its founding in 2022 and values Modular at $1.6 billion – almost tripling our valuation from our last raise. The investment reflects our incredible momentum and reinforces our position as the world’s only truly unified AI infrastructure platform to power the future of AI superintelligence.

Read more on our blog.

Read More
Interviews Tim Davis Interviews Tim Davis

Interview with Alejandro Cremades

An interview with Alejandro Cremades about raising $130 Million to build a next-generation AI Platform to simplify AI development and deployment after multiple startups and acaling AI at Google. You can read the full blog post here, and its embedded below.

Multiple startups, scaling AI at Google, and now raised $130 million to build a next-generation AI platform

Tim Davis's entrepreneurial journey reflects a rare blend of intellectual curiosity, scrappy resilience, and a deep commitment to building relationships and networks. He has an exciting story, which includes an acquihire by Google and the launch of his latest venture, Modular.

Tim’s career trajectory started from gaming on a Commodore 64 and progressed to founding a food delivery startup before Uber Eats. His story isn’t just about pivots and products; it offers valuable lessons in adapting, grit, and navigating ecosystems on both sides of the Pacific.

In a riveting interview on the Dealmakers Podcast, Tim discussed hypergrowth companies, working at a large tech giant like Google, and raising an impressive $130M for Modular. 

Growing Up in Melbourne - The Early Years

Tim was born and raised in Melbourne, Australia, in the 1980s and 1990s. His mom was an artist, and his dad was a banker. Along with his older brother, Tim became an avid gamer, playing and honing his skills at Boulder Dash, Maniac Mansion, and Prince of Persia.

Soon, the brothers were modding many of the games to alter their files and change their behavior, appearance, or even introduce entirely new content. They would also find hacks for the games, programming in BASIC, and eventually, moved to Windows OS, which users preferred in Australia. 

Tim remembers moving quickly to Railway Tycoon and Doom, which fueled his love for technology and computer systems. He also had a keen interest in puzzles and math throughout school. But the path from childhood gamer to Silicon Valley founder wasn’t linear.

Early Curiosity and a Winding Educational Road

Tim’s academic journey is one of the most eclectic you’ll find. He studied chemical engineering and microbiology, then added commerce, mathematics, an MBA, and even a JD to the mix. “I kept flipping a series of interests,” he says.

Tim recalls how he actually ended up finishing specific parts of the course earlier, but lost interest in fluid mechanics, which was a large part of chemical engineering. He progressed to learning actual science and finance, even exploring the possibility of becoming an investment banker. 
But this seemingly scattered path wasn’t aimless; it was a quest for purpose. Despite internships in law firms and banks, Tim found himself disillusioned by the rigid, hierarchical growth trajectories in corporate Australia. 
“Even if you were a rockstar, you had to wait your turn. That felt strange.” The clarity Tim needed came from understanding what he didn’t want–to be stuck in a 9-to-5. “Life’s short,” he says. “While you're young, you’ve got time to explore different things. Why not take a shot?”

Now that he is older and has a family, time has become the most scarce resource available. Tim comments wryly that you don’t have an appreciation for that when you're younger. He was keen on building an interesting business, reasoning that, worst case, he could always lean on his top-tier education. 

As Tim sees it, growing up in Australia gave him the added advantage of a government-backed education, which didn’t saddle him with a massive student debt, unlike in many other countries. 

Entering the Startup Arena: Image Recognition and Fundraising Lessons

Tim’s first foray into startups came while studying patent and trademark law. “I was fascinated by how much effort went into innovation,” he says. Since he had studied computer science and technology, he was inspired to build a business around them. 

That fascination led to an image recognition startup focused on identifying branded content inside photos, an idea rooted in intellectual property logic. Tim began to think about designing a way for brands to find themselves inside the images and monetize them directly.

But the Australian fundraising ecosystem at the time was harsh. “Angels would offer $100K for 20-30% of your company,” Tim explains, a model fundamentally incompatible with scaling. “We realized, if we really want to do this, we need to go to the US.”

Tim saw that eventually they would have to raise several more funding rounds, for the company, CrowdSend, to become profitable. 

But he would have given away a massive percentage to someone who hadn’t fundamentally contributed a considerable amount of capital versus the risk Tim was taking

Moving to the US and Landing in Silicon Valley

In 2012, Tim boarded a one-way flight to Silicon Valley, landing in a hacker house he had found on Airbnb, which was run by a YC founder whose startup had failed. It turned out to be a blessing. The house, a de facto college dorm for ambitious misfits, became the spark for his second startup

At the time, Tim didn’t know much about Silicon Valley, but interacting with the talented folk at the hacker house, he discovered what an incredible place it was. It had lots of entrepreneurs from around the world. 

Arriving in the US, Tim decided against pursuing the image recognition business and instead started a new company with a co-founder he had met, Francisco Magdaleno.

The Early Hustle, Fluc Inc: The Pre-Uber Eats Era Food Delivery Idea

What started as a scrappy food delivery idea among housemates quickly turned into a fully functioning business. Tim built the front end and back end, and his co-founder developed the iOS app. 

Before they knew it, they were pioneering one of the first food delivery apps just as DoorDash was launching in stealth mode as “Palo Alto Delivery.” 

But signing up restaurants was painful. “They laughed at us,” Tim remembers, but he was very confident that their prototype was working well. 

Inspired by Grubhub’s financial statements and business model, which primarily focused on sales and marketing, they added restaurants without permission and increased menu prices by 10% to 20%. The hack worked. Consumers wanted selection. That was the unlock.

The two cofounders continued coding remotely while sorting out their visa situation. They also brought in a third cofounder from the US, Adam Ahmad. When they launched Fluc, it gained popularity quickly since there were no other options in the market that added every restaurant. 

The service exploded in Stanford, Palo Alto, and Mountain View. By 2014, Tim and Francisco were doing millions in top-line revenue. But the margins were brutal. “Food delivery is a horrendous margin business,” Tim admits. 

Back in 2014-2015, the environment was particularly challenging for the on-demand economy. Legal uncertainties about whether drivers were contractors or employees spooked investors, making further fundraising difficult.

Google Steps In: A New Chapter

Amid rising legal complexity and capital challenges, Google entered the picture. The company wasn’t interested in the food delivery business, but they were very interested in the team. Google conducted interviews with the startup’s team members. 

“It wasn’t a formal acquihire,” Tim notes. “They just wanted the people they thought were talented.” Tim and a few others joined Google, while some team members were not selected. What followed was a seven-year stint at one of the most prestigious technology companies.

At the time, Google was scaling a business called Google Express for North American folks. The company was not dissimilar to Instacart, where they were working with merchants to essentially scale the end state delivery.

The Google Years: Culture Shock and Product Execution

Going from startup life to Google was a seismic shift, and Tim picked up important lessons. He developed an understanding of how startups worked, as well as building and assembling things. 

Completing design reviews, product reviews, and engineering reviews, while learning how to create a strong product, was also part of his experience. 

Tim also learned organizational discipline and product execution, gaining valuable exposure to some of the world’s most brilliant minds. The melting pot of diversity and talent density impressed and inspired him.

“In startups, we worked from 7 a.m. to midnight. My first day at Google, people left at 4:30 p.m.,” Tim recalls. While the relaxed pace was jarring, Tim soaked up the best aspects of big tech. After a year in Ads, he moved to Google Brain, the elite AI research unit, which is now Google DeepMind.

It was 2017, before the AI boom, but Tim was hooked. “Being around the world’s best in AI was something you just couldn’t get anywhere else.” Although he had been exposed to recommendation systems inside ads in his logistics startup, deep learning was a new paradigm.

Here, Tim met his future co-founder, Chris Lattner, the creator of the legendary Swift programming language. Tens of millions of developers use this language today, and it drives most of the iOS ecosystem.

The Power of Hyper Networks

Tim emphasizes that the mentorship network he built at Google, particularly at Brain, became one of his most valuable assets. “That core group at Brain, many are now leading the next wave of AI startups like Character.ai, Adept, and others.”

What makes a hyper network? According to Tim, it’s not just about connections, but shared experience. “You work on hard problems with talented people. That trust and track record becomes the foundation for your next venture.”

The Next Chapter: Modular and AI’s Supercycle

Even while at Google back in 2016 to 2018, Tim could see that AI was going to have a massive impact on the world. Much of Google's internal technology was years ahead of the world, and he could see the possibilities. 

Examining the product landscape, Tim noted that NVIDIA owned most of the compute that powers the world’s AI, even in its early stages in 2018-2019. At the time, Google had built its infrastructure called TPUs. 

Tim had a background in product development, marketing, design, and sales, while Chris is a world-renowned engineer. “We asked ourselves—what if there was an open, universal abstraction layer for AI workloads?” Tim explains. 

The vision was a platform where developers could define their model, budget, and latency needs without caring about what hardware ran it, essentially making AI compute truly portable and efficient.

This idea formed the foundation of Modular. It wasn’t a small bet; it was a deep-tech infrastructure play that would take years to build. But the potential to decentralize AI hardware dependency and optimize performance across environments made it one worth pursuing.

Business Model: Scaling with Compute

Modular’s revenue model is tightly tied to usage. “We scale with the amount of compute that flows through our platform,” Tim says. Similar to how Databricks charges based on compute units, Modular’s customers pay based on the volume of AI workloads run.

For enterprises with on-premise deployments, the model shifts to a per-GPU pricing structure. Modular integrates seamlessly with environments like Kubernetes, allowing large-scale AI training or inference across private infrastructure.

Additionally, Modular has embraced cloud partnerships as a key growth channel. Working with providers to embed Modular into their offerings allows the company to monetize through distribution partnerships.

Tim likens this approach to the early Microsoft-Intel alliance or Databricks' partnership with Azure. As he sees it, channel partnerships are an excellent way to get strong distribution. It’s an exciting area that they have also been utilizing. 

A Different Approach to Fundraising

With over $130M raised, including backing from Google, Tim has gained a new perspective on how to raise capital effectively. His key takeaway? Skip the pitch deck and start with a memo, explaining why it was a significant opportunity for the world.

Storytelling is everything that Tim Davis was able to master. The key is capturing the essence of what you are doing in 15 to 20 slides. For a winning deck, take a look at the pitch deck template created by Peter Thiel, Silicon Valley legend (see it here) https://startupfundraising.com/pitch-guide  where the most critical slides are highlighted.

“In our seed round, we didn’t use slides,” Tim explains. “We wrote a three-to-four-page memo explaining what made this opportunity and us as founders uniquely compelling.” The Amazon-style narrative was met with enthusiasm from investors, who appreciated the clarity and depth.

This written approach also enabled more meaningful conversations. “Instead of starting from scratch in meetings, investors came in prepared with thoughtful questions. We’d go straight to whiteboarding,” Tim says. 

The contrast to his first fundraising experience, where pitch decks led to polite but empty rejections, was stark. The memo strategy wasn’t a one-off. Tim and Chris used it again for Modular’s $100M round.

They supplemented it with detailed papers on AI trends, compute economics, and developer growth. The result? Describing things in written form led to deeper discussions and faster alignment with the right investors.

Writing as a Cultural Backbone

Tim strongly encourages people to write their plans in a two-page memo, describing everything they think they can do, why they are uniquely positioned to do it, and why someone should give them the capital over all the competitors.

A two-page memo is more visionary and proves that the project deserves backing. It pushes founders to think deeply since they have only two pages to raise capital on. It also helps gain clarity on what they are building,

Modular’s documentation-first philosophy isn’t just for fundraising. Tim has embedded it into the company’s operating cadence. “We ask everyone to write down decisions using a simple problem-solving framework,” he shares.

The framework is composed of five core questions:

  1. What problem are we solving?

  2. Why is now the right time?

  3. What does success look like?

  4. What alternatives have been considered?

  5. What’s the recommended course of action?

“It’s amazing how this clarity either makes a path obvious or sparks a productive debate,” Tim explains. Whether it’s a hiring decision, a product strategy, or a sales play, the same structure applies. This culture of structured thinking has become central to Modular’s execution engine.

Tim says that he just wants to see a document that briefly outlines the simple architecture. It’s a succinct framework to help drive organizational and product decision-making across companies.

On Mentorship, Networks, and Helping Others

Tim is deeply aware of how relationships have shaped his journey from crashing at a hacker house in his early days to meeting his Modular co-founder at Google Brain. Now, he tries to pay that forward.

“I came to the U.S. from Australia not knowing anyone,” Tim says. “People bet on me, and everything meaningful in Silicon Valley is about how you treat people and the relationships you build.” For founders navigating similar transitions, Tim is generous with his time. 

“If I can help someone get a leg up, I try to. I know how hard it is to land in a new country and try to build something important.”

Final Reflections

Tim Davis’ story is one of navigation across continents, careers, and paradigms. From early missteps in chemical engineering to startup chaos and enterprise calm at Google, he’s assembled a unique blend of technical fluency, legal insight, and network leverage. 

What ties it all together is a founder’s mindset: curious, bold, and constantly evolving. As Tim puts it, “You can’t always plan the path, but you can keep showing up, keep building, and make sure you’re surrounded by the best.”

After years at Google and a successful startup exit, Tim Davis wasn’t just interested in launching another venture; he aimed to reshape how AI workloads are deployed and scaled fundamentally. 

Together with Chris Lattner, the legendary engineer behind the Swift programming language, Tim co-founded Modular, a company building what they see as a missing layer in the AI ecosystem: a hardware-agnostic infrastructure for machine learning workloads.

Listen https://alejandrocremades.com/tim-davis/ to the full podcast episode to know more, including: 

  • Tim Davis's unconventional education and early passion for gaming laid the foundation for a bold entrepreneurial path across industries and continents.

  • His first startup experience taught him the hard lessons of equity, risk, and the limitations of the Australian fundraising ecosystem.

  • Moving to Silicon Valley transformed his trajectory, exposing him to global talent, scrappy startup culture, and ultimately Google’s scale.

  • Google Brain became a pivotal experience, where Tim gained deep exposure to AI and formed critical relationships that fueled his next venture.

  • With Modular, Tim is building a hardware-agnostic platform for AI compute, aiming to decentralize and optimize AI infrastructure.

  • His fundraising strategy, centered on narrative memos instead of pitch decks, has helped raise over $130M and foster deeper investor alignment.

  • Tim’s focus on structured thinking, writing culture, and mentorship reflects his belief in clarity, networks, and paying it forward.

Read More
Essays Tim Davis Essays Tim Davis

Scale or Surrender: When watts determine freedom

Consider this provocative framing: what if we viewed our collective future not through the lens of human populations and national borders, but through available compute capacity? In this view, the race to build massive datacenter infrastructure becomes humanity's defining competition. This perspective makes efficiency not just important, but existential.

Over the past two centuries, humanity's relationship with energy has been nothing short of transformative. If you chart global primary energy consumption from the Industrial Revolution to today, you'll see something remarkable: an almost unbroken ascent, punctuated by only three brief pauses - the early 1980s oil crisis aftermath, the 2009 financial crisis, and the 2020 pandemic. Otherwise, it's been an extraordinary march upward, powered first by coal and oil, then natural gas, nuclear, hydropower, and increasingly, renewables. This wonderful graphic highlights this well - with populous nations like China, the United States, and India dominating total consumption on a per-person basis.

The geographic distribution of this energy consumption tells a striking story. China, the United States, and India dominate in absolute terms, but the per-capita numbers reveal something more profound. Citizens of Iceland, Norway, Canada, the United States, and wealthy Gulf states like Qatar and Saudi Arabia consume up to 100 times more energy than those in the world's poorest regions. This isn't merely inequality - it's a chasm so vast that millions of people still rely on traditional biomass (wood, agricultural residues) that doesn't even register in our global energy statistics, creating data gaps

The disparities in electricity generation are equally stark. Iceland, blessed with abundant geothermal and hydro resources, generates hundreds of times more electricity per person than many low-income nations, where annual per-capita generation can fall below 100 kilowatt-hours - less than what a modern refrigerator uses in two months.

This context matters immensely as we confront the dual challenge of our time: meeting rising global energy demand while urgently decarbonizing our energy supply. Despite record investments in clean technologies, fossil fuels still account for approximately 81.5% of global primary energy. The math here is unforgiving - renewable sources must not only meet all new demand but also replace existing fossil fuel capacity if we're to bend the emissions curve downward.

Enter artificial intelligence, with its voracious and growing appetite for electricity.

In 2023, U.S. data centers consumed approximately 176 terawatt-hours - 4.4% of national electricity consumption. Current projections suggest this could reach 325 to 580 TWh by 2028, representing 6.7% to 12% of total U.S. electricity demand driven largely by AI workloads that demand ever-increasing compute power and specialized hardware. To contextualize these numbers: we're talking about enough electricity to power between 32.5 and 58 million American homes.

The AI industry has long understood a critical metric that deserves wider attention: tokens-per-dollar-per-watt. This measure of computational efficiency relative to both cost and energy consumption has been a focus at Google and other leading technology companies for years. It represents the kind of systems thinking we desperately need as AI capabilities expand.

The challenge before us is clear. We're attempting to build transformative AI systems while simultaneously addressing the climate crisis. These goals aren't inherently incompatible, but reconciling them requires unprecedented coordination and innovation across multiple domains:

  • Hardware efficiency: Next-generation chips that deliver dramatically better performance-per-watt

  • Operational intelligence: Carbon-aware scheduling that aligns compute-intensive tasks with renewable energy availability

  • Infrastructure innovation: On-site renewable generation and novel cooling systems that minimize overhead

  • System integration: Data centers that contribute to local energy systems through waste heat recovery

  • Radical transparency: Clear reporting standards that drive competition on efficiency metrics

Global energy consumption tells a story of both peril and promise. As artificial intelligence scales exponentially, it threatens to derail climate progress - yet history shows us that human ingenuity consistently reimagines our energy systems when survival demands it. We have already proven we can build transformative AI; the defining challenge now is whether we can build it sustainably, ensuring our creations enhance rather than endanger the world they serve.

The stakes are higher than they appear. Even breakthrough efficiency gains in AI hardware may paradoxically increase total energy consumption - a manifestation of Jevon's Paradox, where technological improvements drive greater overall demand. At this crossroads of intelligence and energy transformation, our choices will determine whether AI becomes humanity's greatest tool or its most consequential miscalculation.

The arithmetic is challenging, but not impossible. What's required is the kind of systematic thinking and ambitious action that has characterized humanity's greatest technological leaps. The alternative - allowing AI's energy demands to grow unchecked - would represent a profound failure of imagination and responsibility, but also the risks are enormous as whichever nations control the most powerful AI systems - are the new superpowers of tomorrow. In this article, I try to shine a light on what's causing the enormous growth of energy demands and some thoughts about the path forward.

The geography of American power

To truly grasp the magnitude of AI's growing energy demands, it's instructive to examine America's electricity generation landscape. At the apex sits the Palo Verde Nuclear Generating Station in Arizona, the nation's largest power producer, generating approximately 32 million megawatt-hours annually - equivalent to 32 billion kWh or 32 TWh.

What does 32 billion kWh actually mean? The U.S. Energy Information Administration reports that the average American household consumes about 10,500 kWh per year. Simple arithmetic reveals that Palo Verde alone could theoretically power 3.05 million homes - roughly 2.5% of the nation's 120.92 million households. One facility, powering the equivalent of a major metropolitan area.

The roster of America's electricity giants tells a fascinating story about our energy infrastructure. After Palo Verde, we have Browns Ferry (31 TWh, nuclear), Peach Bottom (22 TWh, nuclear), and then Grand Coulee Dam (21 TWh) - the hydroelectric marvel that helped build the American West. The list continues with West County Energy Center (19 TWh, natural gas), W.A. Parish (16 TWh, a coal/gas hybrid), and Plant Scherer in Georgia (15 TWh, coal).

Notice the pattern? Nuclear dominates the top tier, followed by a mix of hydro, gas, and coal. After these giants, capacity drops precipitously to facilities generating around 3 GWh - a reminder of how concentrated our electricity production really is. This concentration matters. When we project data centers consuming 325-580 TWh by 2028, we're talking about the equivalent of 10-18 Palo Verde stations running exclusively to power AI and digital infrastructure. That's not replacing existing demand - that's additional load on a grid already straining to decarbonize.

The average U.S. household consumes about 10,500 kilowatthours (kWh) of electricity per year, though this varies significantly by region and housing type. Residential electricity primarily powers essential systems: space cooling, water heating, space heating, along with refrigeration, lighting, and electronics. Commercial buildings have vastly different consumption patterns depending on their size and type, ranging from small offices to large retail centers and office complexes, each with varying HVAC, lighting, and operational equipment needs.

The sectoral breakdown of U.S. electricity consumption reveals a more balanced distribution than commonly understood. According to EIA forecast 2025-2026 power sales will rise to 1,494 billion kWh for residential consumers, 1,420 billion kWh for commercial customers and 1,026 billion kWh for industrial customers with longer term forecasts still mostly within norms. This translates to approximately 38% residential, 36% commercial, and 26% industrial consumption. Rather than the residential sector being overshadowed by commercial and industrial users, it actually represents the largest single sector of electricity demand, with commercial consumption running a close second. This distribution reflects America's transition toward greater electrification in homes and businesses, driven by factors including growing demand from artificial intelligence and data centers and as homes and businesses use more electricity.

The long view is revealing. According to the Energy Information Administration, U.S. electricity consumption increased in all but 11 years between 1950 and 2022. The rare declines - including 2019, 2020, and 2023 - coincided with economic contractions, efficiency improvements, or exceptional circumstances like the pandemic. The overarching trend remains unmistakably upward. While the exact figures here may carry some uncertainty, they accurately capture the essential dynamics. What matters isn't whether commercial usage is precisely 6.7 times residential, but that the disparity is substantial and the growth trajectory is clear. These patterns - concentrated commercial demand and relentless growth - form the backdrop against which we must evaluate AI's emerging energy requirements.

Understanding these scales helps frame the challenge ahead. Every percentage point of national electricity consumption that shifts to data centers represents millions of homes' worth of power. The infrastructure required to meet this demand sustainably doesn't just appear - it must be planned, financed, and built, all while racing against both growing demand and climate imperatives.

The numbers that keep me up at night

Let's revisit the core projection: U.S. data centers consumed approximately 176 TWh in 2023 (4.4% of national electricity) and are projected to reach 325-580 TWh by 2028 – equivalent to powering 32.5 to 58 million American homes.

But here's what keeps me up at night: 580 TWh might be just the beginning of what we need.

Consider today's reality. One analysis estimated that ChatGPT inference alone consumes an estimated 226.8 GWh annually – enough to power 21,000 U.S. homes – and that's already outdated. The International Energy Association (IEA) offers a more sobering projection: 945 TWh - that's the entire electricity consumption of the world's third-largest economy - Japan. 

Let that sink in - AI could require as much electricity as the world's third-largest economy. The composition of this demand has already shifted fundamentally. During my time at Google, I watched inference overtake training as the primary driver of compute demand through the late 2010s, and now inference is quickly rising to represent more than 80%+ of the AI compute capacity across the industry. This matters because while training happens in discrete, intensive bursts, inference runs continuously at scale, serving billions of requests around the clock - every query, every recommendation, every generated response adds to the load. This is the crux of our challenge: AI follows exponential growth patterns that surprise even those who've spent years watching them unfold. Given the convergence toward dominant model architectures, inference's share will likely climb further, meaning our upper-bound projection of 945 TWh could see inference alone consuming over 756 TWh by 2028.

But doesn’t edge computing promise to slash data center demands? This is a narrative I've heard repeatedly throughout years of scaling early edge AI systems. Yet I remain deeply skeptical of any order-of-magnitude impact. The reason is simple: we've barely scratched the surface of enterprise, government, and industrial AI adoption. These sectors will unleash computational demands that dwarf any efficiency gains from consumer devices processing locally. Consider the asymmetry: for every smartphone performing local voice recognition, hundreds of enterprise systems are analyzing documents, monitoring infrastructure, processing surveillance footage, and generating complex reports. The sheer scale of this institutional transformation will eclipse whatever load we shift to the edge.

This reality leads us to the heart of the matter: if inference drives our energy challenge, how do we understand its consumption patterns? What does the energy anatomy of inference reveal, and where might we find our leverage points for optimization? Understanding these patterns isn't just an academic exercise - it's essential for developing strategies that can accommodate AI's growth while continuing to grow our energy infrastructure. The always-on nature of inference, combined with its direct relationship to usage, creates a fundamentally different challenge than the periodic spikes of model training.

Inference: the GOAT of consumption

Inference footprint represents the electricity consumed each time an AI model generates a response - as AI becomes ubiquitous across digital services, inference will inevitably dominate long-term energy costs. This raises a crucial question: how do we properly measure inference energy consumption? What's the right framework for calculating inference-per-token-per-watt?

Let's develop a working model, with an important caveat: these calculations rest on rough assumptions about the current AI landscape. They presume most AI continues running on transformer architectures without fundamental changes over the next few years - though I suspect this assumption may prove conservative. We're likely to use AI itself to discover more efficient architectures, potentially invalidating these projections in favorable ways. With that context, let's examine how inference energy consumption actually works and what drives its costs at scale.

The quadratic curse

Transformers are the neural network architecture powering most modern AI systems - from ChatGPT to Claude to Gemini, created by former colleagues at Google. The key innovation of this architecture is the ability to process all parts of an input simultaneously while understanding relationships between distant elements in the text. In transformer inference, prefill is the initial computational phase where the model processes your entire input prompt before generating any output. This involves a single forward pass through the network, computing hidden representations for all input tokens at once.

Your sequence length simply counts these tokens - the basic units of text that might be letters, partial words, or whole words depending on the tokenizer. "Hello, world!" typically translates to 3-4 tokens, while a lengthy document might contain thousands. This distinction matters because prefill computation scales with sequence length, making long prompts significantly more energy-intensive than short ones.

Prefill time grows quadratically with sequence length - double the input, quadruple the computation. This scaling behavior stems from transformers' core mechanism: self-attention. Self-attention requires computing relationships between every pair of tokens in the input. For n tokens, that's comparisons. Unlike older architectures (RNNs) that process tokens sequentially, transformers examine all tokens simultaneously, with each token gathering information from every other token in parallel.

Here's an intuitive analogy: imagine a roundtable discussion where each participant (token) prepares three items:

  • Query: "What information am I seeking?"

  • Key: "What information do I possess?"

  • Value: "What insight can I contribute?"

Each participant shares their query with everyone else, comparing it against others' keys to find the most relevant matches. They then synthesize their understanding by combining values from those whose keys best align with their query. Every participant does this simultaneously, creating a rich, interconnected understanding of the entire conversation. This elegant mechanism enables transformers' remarkable capabilities, but it comes at a cost: computational requirements that scale quadratically with input length. A 2,000-token prompt requires four times the computation of a 1,000-token prompt, not twice. This mathematical reality shapes the energy economics of AI inference at scale.

The Two Phases of Transformer Processing

Every transformer request involves two distinct computational phases:

  1. Prefill: Processing the entire input prompt (quadratic scaling with input length, O(n²) complexity)

  2. Decode: Generating output tokens one by one (linear scaling with output length, O(n) complexity)

This scaling difference has profound implications. While decode time grows linearly with the number of tokens generated, prefill time grows with the square of input length. The longer your prompt, the more dramatically prefill dominates total processing time.

Consider the relative computational work (in arbitrary units, assuming 50 output tokens):

Input Length Prefill Work (∝ n²) Decode Work Prefill % of Total
500 tokens 250,000 units 50,000 units 83%
1,000 tokens 1,000,000 units 50,000 units 95%
2,000 tokens 4,000,000 units 50,000 units 99%

The pattern is stark. At 500 input tokens, prefill already consumes 83% of processing time. Double the input to 1,000 tokens, and prefill jumps to 95% - the actual generation phase becomes almost negligible. At 2,000 tokens, you're spending 99% of compute just understanding the prompt. Here's what's happening:

  • Prefill work = n² (where n = input tokens)

  • Decode work = 1,000 × output tokens (arbitrary scaling factor)

This quadratic scaling of self-attention explains why long-context models are so computationally expensive. As context windows expand from thousands to hundreds of thousands of tokens, the energy requirements don't just grow - they explode. Understanding this dynamic is crucial for anyone designing AI systems or planning infrastructure for the age of ubiquitous AI.

Doubling words, quadrupling watts

The quadratic scaling of context windows isn't just an abstract computational concern - it translates directly into energy consumption. Every FLOP requires energy, and when FLOPs scale quadratically, so does your electricity bill.

The energy equation is straightforward:

  • Devices draw roughly constant power P during operation (e.g., 300W for a high-performance GPU)

  • Energy consumed equals power multiplied by time: E = P × T

  • Since prefill time scales quadratically with input length, so does prefill energy

Let's make this concrete with realistic parameters:

  • Power draw: 300W

  • Decode time: 20ms (fixed for 50 output tokens)

  • Baseline prefill: 100ms for 500 input tokens

Input Tokens Prefill Time Prefill Energy Decode Time Decode Energy Prefill % of Total
500 0.10 s 30 J 0.02 s 6 J 83%
1,000 0.40 s 120 J 0.02 s 6 J 95%
2,000 1.60 s 480 J 0.02 s 6 J 99%

The energy story mirrors the computational one. While decode energy remains constant at 6 joules regardless of input length, prefill energy explodes from 30J to 480J as input doubles from 500 to 2,000 tokens. At 2,000 tokens, you're burning 80 times more energy understanding the prompt than generating the response. 

Let's recap these results.

At 500 input tokens, prefill consumes 30J versus decode's 6J - already 83% of total energy. Double the input to 1,000 tokens, and prefill time quadruples, pushing energy consumption to 120J and commanding 95% of the total. By 2,000 tokens, the imbalance becomes extreme: 480J for prefill versus 6J for decode, with prefill consuming 99% of the energy budget. Extrapolate to a 10,000-token prompt generating just 1500 output tokens, and you're looking at 3.4Wh per query - nearly all spent on prefill. This isn't a marginal effect; it's the dominant factor in inference energy consumption. 

The implications here are therefore profound. Whether you're designing for on-device inference with battery constraints, deploying in autonomous vehicles, or managing massive cloud infrastructure costs at scale - prompt length becomes your primary lever for controlling energy consumption. The quadratic scaling means that doubling prompt length doesn't double energy use - it roughly quadruples it. This scaling asymmetry defines the energy economics of AI. Decode plods along linearly - each output token costs the same as the last. But prefill explodes quadratically with input length, its hunger growing with the square of every token fed. A thousand-token prompt doesn't just double the cost of a 500-token prompt - it quadruples it. By the time we reach today's massive contexts, decode disappears entirely, a thin shadow cast by prefill's towering consumption.

And yet, users control neither the true input nor output. There isn’t a “token budget” in any service out there today, and that would likely create a frustrating user experience if there was. The largest providers - OpenAI, Google, Anthropic - inject substantial hidden context into every prompt while keeping their system instructions opaque. Output remains equally unconstrained: unless users explicitly demand brevity, models generate tokens freely and most services never think to limit responses, and most users don’t even understand they can.

This creates a fundamental tension in AI system design. While longer contexts enable richer interactions and more sophisticated reasoning, they exact an exponentially increasing energy toll. Once prompts exceed a few hundred tokens, virtually all computational resources are consumed by the prefill phase alone. For sustainable AI deployment at scale, prompt concision isn't merely good practice - it's an energy imperative. The difference between 500-token and 2,000-token average prompts could determine whether our global infrastructure remains viable or collapses under its own consumption.

This problem compounds as AI agents and capabilities like Deep Research proliferate. Each autonomous action, each recursive query, each unconstrained generation adds to an already exponential curve. We're building systems designed to think deeply while hoping they'll somehow learn restraint—a contradiction that grows more stark with every token generated.

The paradox blooms in plain sight

Here's the irony: while physics demands shorter prompts, the industry is sprinting in the opposite direction. Claude 4's system prompt far exceeds 10,000 tokens. Despite optimization techniques like KV cache retrieval and prefix caching, the overwhelming trend is toward ever-expanding context windows. We're stuffing everything we can into prompts - documentation, code repositories, conversation histories - because it demonstrably improves model capabilities.

A colleague recently quipped (hat tip, Tyler!): "We'll achieve AGI when all of Wikipedia fits in the prompt!" It's a joke that hits uncomfortably close to our current trajectory of maximizing context windows at every opportunity. 

My rough calculations above turn out to align remarkably well with recent empirical findings. This very recent research shows that a GPT-4o query with 10,000 input tokens and 1,500 output tokens consumes approximately 1.7 Wh on commercial datacenter hardware. For models with more intensive reasoning capabilities, the numbers climb dramatically: DeepSeek-R1 averages 33 Wh for long prompts, while OpenAI's o3 model reaches 39 Wh. These aren't theoretical projections - they're measured consumption figures from production systems. 

The energy cost of our context window expansion is real, substantial, and growing with each new model generation. We're caught between two competing imperatives: the computational benefits of longer contexts and the exponential energy costs they incur. The other interesting observation on this table below, is the explosive increase in computational power requirements as models have gotten larger and more sophisticated - the power laws of exponential scaling continue.

Model Release Date Energy Consumption (Wh)
(100 input-300 output tokens)
Energy Consumption (Wh)
(1K input-1K output tokens)
Energy Consumption (Wh)
(10K input-1.5K output tokens)
o4-mini (high)Apr 16, 20252.916 ± 1.6055.039 ± 2.7645.666 ± 2.118
o3Apr 16, 20257.026 ± 3.66321.414 ± 14.27339.223 ± 20.317
GPT-4.1Apr 14, 20250.918 ± 0.4982.513 ± 1.2864.233 ± 1.968
GPT-4.1 miniApr 14, 20250.421 ± 0.1970.847 ± 0.3791.590 ± 0.801
GPT-4.1 nanoApr 14, 20250.103 ± 0.0370.271 ± 0.0870.454 ± 0.208
GPT-4o (Mar '25)Mar 25, 20250.421 ± 0.1271.214 ± 0.3911.788 ± 0.363
GPT-4.5Feb 27, 20256.723 ± 1.20720.500 ± 3.82130.495 ± 5.424
Claude-3.7 SonnetFeb 24, 20250.836 ± 0.1022.781 ± 0.2775.518 ± 0.751
Claude-3.7 Sonnet ETFeb 24, 20253.490 ± 0.3045.683 ± 0.50817.045 ± 4.400
o3-mini (high)Jan 31, 20252.319 ± 0.6705.128 ± 1.5994.596 ± 1.453
o3-miniJan 31, 20250.850 ± 0.3362.447 ± 0.9432.920 ± 0.684
DeepSeek-R1Jan 20, 202523.815 ± 2.16029.000 ± 3.06933.634 ± 3.798
DeepSeek-V3Dec 26, 20243.514 ± 0.4829.129 ± 1.29413.838 ± 1.797
LLaMA-3.3 70BDec 6, 20240.247 ± 0.0320.857 ± 0.1131.646 ± 0.220
o1Dec 5, 20244.446 ± 1.77912.100 ± 3.92217.486 ± 7.701
o1-miniDec 5, 20240.631 ± 0.2051.598 ± 0.5283.605 ± 0.904
LLaMA-3.2 1BSep 25, 20240.070 ± 0.0110.218 ± 0.0350.342 ± 0.056
LLaMA-3.2 3BSep 25, 20240.115 ± 0.0190.377 ± 0.0660.573 ± 0.098
LLaMA-3.2-vision 11BSep 25, 20240.071 ± 0.0110.214 ± 0.0330.938 ± 0.163
LLaMA-3.2-vision 90BSep 25, 20241.077 ± 0.0963.447 ± 0.3025.470 ± 0.493
LLaMA-3.1-8BJul 23, 20240.103 ± 0.0160.329 ± 0.0510.603 ± 0.094
LLaMA-3.1-70BJul 23, 20241.101 ± 0.1323.558 ± 0.42311.628 ± 1.385
LLaMA-3.1-405BJul 23, 20241.991 ± 0.3156.911 ± 0.76920.757 ± 1.796
GPT-4o miniJul 18, 20240.421 ± 0.0821.418 ± 0.3322.106 ± 0.477
LLaMA-3-8BApr 18, 20240.092 ± 0.0140.289 ± 0.045
LLaMA-3-70BApr 18, 20240.636 ± 0.0802.105 ± 0.255
GPT-4 TurboNov 6, 20231.656 ± 0.3896.758 ± 2.9289.726 ± 2.686
GPT-4Mar 14, 20231.978 ± 0.4196.512 ± 1.501

Table 4: How Hungry is AI? , I added Model Release Dates

Trade 8,500 conversations, to keep your home cool

To grasp the practical implications, consider this sobering calculation: at o3's extreme consumption of 39 Wh per query, approximately 76,923 interactions would drain 3,000 kWh - equivalent to powering a typical American home's air conditioning for an entire year. But as users inevitably gravitate toward richer prompts - say, 30K input tokens with 3.5K outputs - the quadratic curse strikes with mathematical precision. Prefill energy multiplies ninefold, collapsing that annual budget to just 8,500 interactions: merely 23 queries per day. This comparison transforms abstract energy figures into visceral reality, revealing how what appears negligible at the per-query level becomes a massive aggregate demand. And we're still only discussing text models.

The trajectory becomes even more stark when we consider multimodal AI. Image and video models can routinely process upwards of hundreds of thousands of tokens per query. Each frame, each visual element, each temporal relationship adds to the token count. As these models become mainstream, we're not just scaling linearly with adoption - we're multiplying adoption rates by dramatically higher per-query energy costs.

The math is unforgiving: widespread deployment of long-context AI at current efficiency levels would require energy infrastructure on a scale we're not remotely prepared for. This isn't a distant concern - it's the reality we're building toward with every context window expansion and every new multimodal capability. The urgency is real, we need to move faster or soon the trade-offs become explicit: 8,500 conversations or one cool home. Your ChatGPT query or your neighbor's heating. And while you might laugh, it’s already happening here in the United States, and water availability is next.

The efficiency mirage, better hardware alone isn’t enough

One way to conclude all of this is to say “But wait, the hyper scalers like Google and others aren’t buying Nuclear Power plants? Surely, we’ll be ok?” I would counter that by firstly saying, actually they are and secondly the vast majority of the world isn’t as sophisticated as the hyperscalers. Google represents the exception, not the rule, in AI efficiency. That's to say, that rather than accepting quadratic scaling as inevitable, Google has:

  • Implemented multiple optimizations that compound together (e.g. software and hardware co-design with TPUs)

  • Focused on inference efficiency where the bulk of tokens are processed

  • Heavy research investment to continue to find alternatives beyond pure transformer architectures where appropriate

  • Achieved order of magnitude efficiency gains that largely offset quadratic scaling as a result of all of these compounding

During my time at Google, the company's sophistication in AI infrastructure was staggering. Through software-hardware co-design with TPUs, relentless focus on inference optimization, and architectural innovations beyond pure transformers, Google achieved order-of-magnitude efficiency gains that largely offset quadratic scaling. This luxury—born from inventing the transformer architecture itself - remains unavailable to most players, including even technology giants like Microsoft and Meta, who are trying to copy the success of the TPU and the most advanced model companies like OpenAI, Anthropic who are trying to develop own custom hardware today

Beyond these elite players lies a wasteland of inefficiency. Industry-wide GPU utilization averages a shocking 15-50%, despite NVIDIA and AMD claiming 90%+ efficiency is achievable. Microsoft's study of 400 real deep learning deployments confirms this reality: enterprise GPU utilization rarely exceeds 50%. Even Meta's Llama 3 405B, running on 16,384 H100 GPUs, achieves only 38% Model Flop Utilization. The physics compounds the problem. NVIDIAs H100 consumes 700W at peak, the A100 draws 400-500W and AMD's MI300X reaches 750W - yet at 25% utilization, these still draw 35-43% of maximum power due to static components like memory controllers. This non-linear power curve creates a cruel efficiency trap: most organizations operate in the steepest part of the curve, where marginal performance gains demand exponential energy increases.

GPU Power Consumption: Comparing Non-Linear Curves

How NVIDIA H100, A100, and AMD MI300X consume power at different utilization levels

NVIDIA H100 700W Peak
Idle Power: 80W
Power @ 15%: 245W (35%)
Power @ 50%: 420W (60%)
Efficiency Loss @ 15%: 133%
NVIDIA A100 500W Peak
Idle Power: 60W
Power @ 15%: 175W (35%)
Power @ 50%: 300W (60%)
Efficiency Loss @ 15%: 133%
AMD MI300X 750W Peak
Idle Power: 85W
Power @ 15%: 263W (35%)
Power @ 50%: 450W (60%)
Efficiency Loss @ 15%: 133%

Key Utilization Points Comparison

Utilization H100 (700W) A100 (500W) MI300X (750W) Power Efficiency
0% (Kernel Idle) 16W 12W 17W ~2-3% of peak
15% (Typical) 245W 175W 263W 35% power for 15% work
50% (Moderate) 420W 300W 450W 60% power for 50% work
70% (Optimal) 560W 400W 600W 80% power for 70% work
100% (Peak) 700W 500W 750W 100% power for 100% work

Lets bring it all back to our token-per-watt metric - my point here is that performance efficiency metrics reveal stark contrasts between theoretical capabilities and real-world achievements. Consider the stark gap between marketing and reality. The H100 delivers an impressive 4.3-5.7 tokens per watt in optimized configurations. But at typical 15% utilization, this plummets to 0.65-0.86 tokens per watt - an 85% efficiency collapse. No amount of hardware innovation can overcome fundamental deployment incompetence. All this highlights my point - software is just as important to energy efficiency as the hardware innovation itself and if you aren’t holding it right - you’re token-per-watt plummets irrespective of the power of the hardware.

The broader implications are staggering. For example, with 3.76 million datacenter GPUs sold in 2023 operating at 15% average utilization, the industry wastes $12.6 billion in underutilized capacity annually, generates 94 million tons of CO2 equivalent to 20.4 million cars, and squanders enough electricity to power 1.1 million American homes. According to SemiAnalysis, each H100 server, carrying a $106,752 annual total cost of ownership, effectively costs organizations $7,413 per useful GPU-month at these 15% utilization levels instead of $1,235 at proper utilization like 90%.

The tools exist to solve this crisis - GPU utilization monitoring, workload optimization, multi-instance allocation - but the industry lacks the sophistication to deploy them effectively. Outside a handful of elite players, the AI revolution runs on fundamentally broken economics -  in regular conversations with many sophisticated technology enterprises - they are closer to 50% utilization for their GPUs, than 90%. Until we close this efficiency gap, hardware improvements will only enable more waste at larger scales, making energy consumption our defining constraint rather than computational capability.

The stakes have never been higher

The mathematics reveals an unforgiving truth: prefill computation scales quadratically with input length, meaning each doubled prompt quadruples energy consumption. As models chase ever-expanding context windows, this fundamental relationship drives AI's steepening energy curve. Meanwhile, the industry's abysmal GPU utilization - averaging 15-50% - compounds the crisis through pure waste. We face a perfect storm: exponentially growing computational demands colliding with systematic inefficiency.

Of course, these projections assume a degree of technological stasis. Innovation could disrupt these trends - and likely will - and the work we are doing at Modular is certainly trying to help. But consider this provocative framing: what if we viewed our collective future not through the lens of human populations and national borders, but through available compute capacity? In this view, the race to build massive datacenter infrastructure becomes humanity's defining competition. If each AI agent represents some fraction of human productive capacity, then the first to achieve a combined human-digital population of 5 or 10 billion wins the AI race, and invents the next great technological frontier.

This perspective makes efficiency not just important, but existential. The promise of AI as humanity's great equalizer inverts into its opposite: a world where computational capacity becomes the new axis of dominance. Without breakthrough innovation at every layer - from silicon to algorithms to system architecture - we face an AI revolution strangled by the very physics of power generation. The future belongs not to the wise, but to the watt-rich. And so we must scale with unprecedented urgency - nuclear, renewables, whatever it takes - because the stakes transcend mere technological supremacy. In this race, computational power becomes political power, and China is currently winning by a large margin. If democratic nations cede the AI frontier to autocracies, we don't just lose a technological edge; we risk watching the values of human dignity and freedom dim under the shadow of algorithmic authoritarianism. The grid we build today determines whose values shape the world of tomorrow.

It's interesting to reflect that we are teaching machines to think with the very energy that makes our planet uninhabitable - yet these same machines may be our only hope of learning to live within our means. AI is both the fever and the cure, the flood and the ark, the hunger that's outrunning itself. We race against our own creation, betting that the intelligence we birth from burning carbon will show us how to stop burning it altogether. The question of our age: Can we make AI wise enough to save us before it grows hungry enough to consume us? The stakes are incredibly unquestionably high in the race to AI superintelligence.

I'll close with an irony that perfectly captures the current moment - the suggestions below come courtesy of Anthropic’s Claude 4 Opus:

  1. Better Chips and Smarter Cooling - The latest AI chips use way less energy for the same work. Pair that with innovative cooling like liquid systems or modular designs, and data centers have already seen energy savings of up to 37% in test runs.

  2. Timing is Everything - Not all AI work needs to happen right now. By running non-urgent tasks when electricity is cleaner (like when it's sunny or windy), some companies have cut their carbon emissions from AI jobs by 80-90%. It's like doing laundry at night when rates are lower, but for the planet.

  3. Power Where You Need It - Building solar panels, wind turbines, and battery storage right at data center sites makes sense. Google's recent $20 billion investment in clean energy shows how tech giants can grow their AI capabilities without relying entirely on the traditional power grid.

  4. Working Together on Infrastructure - Data center operators need to share their growth plans with utility companies. This helps everyone prepare for the massive power needs coming to tech hubs like Northern Virginia, Texas, and Silicon Valley - think of it as giving the power company a heads-up before throwing a huge party.

  5. Show Your Work - Just like appliances have energy ratings, AI companies should tell us how much power they're using. Whether its energy per query or per training session, transparency creates healthy competition to be more efficient.

  6. Investing in Tomorrow's Solutions - Government programs are funding research into game-changing technologies like optical processors that could use 10 times less energy. There's also exciting work on making AI models smaller and smarter without losing capabilities.

  7. Turning Waste into Resources - Data centers generate tons of heat - why not use it? Some facilities are already warming nearby buildings with their excess heat, turning what was waste into a community benefit.

And there it is - 40 watt-hours spent asking AI how to save token-per-watt-hours, with some of these ideas still unproven (e.g. optical processors). The perfect metaphor for our moment: we burn the world to ask how to stop burning it, racing our own shadow toward either wisdom or ruin. The verdict seems more absolute: scale or surrender - there is no middle ground in the physics of power.

Thanks to Tyler Kenney, Kalor Lewis, Eric Johsnon, Azeem Azhar, Christopher Kauffman, Will Horyn, Jessica Richman, Duncan Grove and others for many fun discussions on the nature of AI and energy. 

Read More
Interviews Tim Davis Interviews Tim Davis

Fund/Build/Scale Podcast Interview

Leaving a high-paying role at Google to take on NVIDIA, Intel, and AMD is not for the faint of heart, but that’s exactly what Tim Davis, co-founder and president of Modular, did.

I had a wonderful conversion with Walter Thompson about Modular some time ago, we covered AI compute, building a new AI software stack, the ups and downs of startups and scaling AI into the future.

Introduction

Leaving a high-paying role at Google to take on NVIDIA, Intel, and AMD is not for the faint of heart, but that’s exactly what Tim Davis, co-founder and president of Modular, did.

In this episode of Fund/Build/Scale, Tim explains why Modular has raised $130M to reimagine AI compute infrastructure, and what he’s learned trying to build a platform that competes with some of the biggest names in tech.

We talked about:

  • 🚀 Why Modular believes AI workloads need a hardware-agnostic execution platform

  • 💡 How Tim and co-founder Chris Lattner decided to “start from the hardest part of the stack”

  • 💰 The trade-offs of raising VC for infrastructure-heavy startups

  • 🛠 Why Modular focuses on talent density and how they’ve recruited top engineers from Google, NVIDIA, and beyond

  • 🌍 What it takes to break into the AI space when your competitors are trillion-dollar companies

Tim takes a thoughtful, deep-dive approach to this conversation—unpacking the complexities of AI infrastructure and what it takes to build in one of tech’s most competitive spaces. There’s valuable insight here for founders navigating technical markets or aiming to disrupt entrenched players.

Episode Breakdown

(1:26) “We are building a new accelerated execution platform for compute.”

(6:41) “ It will exist all over the place and it already does, but AI will be everywhere that compute is.”

(11:18) “ You only you only have so much time in a week. What is the thing that you're best at?”

(15:13) “ We have decided to start from the hardest part of the software stack.”

(22:44) “For the most talented people in the world, the risk is actually not as great as what you think.”

(30:24) “ Growing up in Australia, my view of the of the United States was very much driven from the media and from Hollywood.”

(33:26) “ I sat in a room for six weeks and just met everyone that I could. And that really was the beginning of a journey to the United States.”

(37:48) “ I still think there's a special place in the Bay Area, and in the United States, there is a different risk appetite.”

(40:41) The one question Tim would have to ask the CEO before he’d take a job at someone else’s early-stage startup.

Read More
Essays Tim Davis Essays Tim Davis

AI Regulation: step with care, and great tact

AI systems take an incredible amount of time to build and get right - I know because I have helped scale some of the largest AI systems in the world, which have directly and indirectly impacted billions of people. If I step back and reflect briefly - we were promised mass production self-driving cars 10+ years ago, and yet we still barely have any autonomous vehicles on the road today. Radiologists haven’t been replaced by AI despite predictions virtually guaranteeing as much, and the best available consumer robot we have is the iRobot J7 household vacuum

“So be sure when you step, Step with care and great tact. And remember that life's A Great Balancing Act. And will you succeed? Yes! You will, indeed! (98 and ¾ percent guaranteed).” ― Dr. Seuss, Oh, the Places You'll Go!

AI systems take an incredible amount of time to build and get right - I know because I have helped scale some of the largest AI systems in the world, which have directly and indirectly impacted billions of people. If I step back and reflect briefly - we were promised mass production self-driving cars 10+ years ago, and yet we still barely have any autonomous vehicles on the road today. Radiologists haven’t been replaced by AI despite predictions virtually guaranteeing as much, and the best available consumer robot we have is the iRobot J7 household vacuum

We often fall far short of the technological exuberance we project into the world, time and time again realizing that producing incredibly robust production systems is always harder than we anticipate. Indeed, Roy Amara made this observation long ago:

We tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run.

Even when we have put seemingly incredible technology out into the world, we often overshoot and release it to minimal demand - the litany of technologies that promised to change the world but subsequently failed is a testament to that. So let's be realistic about AI - we definitely have better recommendation systems, we have better chatbots, we have translation across multiple languages, we can take better photos, we are now all better copywriters, and we have helpful voice assistants. I've been involved in using AI to solve real-world problems like saving the Great Barrier Reef, and enabling people who have lost limbs to rediscover their lives again. Our world is a better place because of such technologies, and they have been developed and deployed into real applications over the last 10+ years with profoundly positive effect.

But none of this is remotely close to AGI - artificial general intelligence - or anywhere near ASI - artificial superintelligence. We have a long way to go despite what the media presages, like a Hollywood blockbuster storyline. In fact, in my opinion, these views are a bet against humanity because they gravely overshadow the incredible positives that AI is already, and will continue, providing - massive improvements in healthcare, climate change, manufacturing, entertainment, transport, education, and more - all of which will bring us closer to understanding who we are as a species. We can debate the merits of longtermism forever, but the world has serious problems we can solve with AI today. Instead, we should be asking ourselves why we are scrambling to stifle innovation and implement naively restrictive regulatory frameworks now, at a time when we are truly still trying to understand how AI works and how it will impact our society? 

Current proposals for regulation seem more concerned with the ideological risks associated with transformative AGI - which, unless we uncover some incredible change in physics, energy, and computing, or perhaps uncover that AGI exists in a vastly different capacity to our understanding of intelligence today - is nowhere near close. The hysteria projected by many does not comport with the reality of where we are today or where we will be anytime soon. It's the classic mistake of inductive reasoning - taking a very specific observed example and making overly broad generalizations and logical jumps. Production AI systems do far more than just execute an AI model - they are complex systems - similarly, a radiologist does far more than look at an image - they treat a person.

For AI to change the world, one cannot point solely to an algorithm or a graph of weights and state the job is done - AI must exist within a successful product that society consumes at scale - distribution matters a lot. And the reality today is that the AI infrastructure we currently possess was developed primarily by researchers for research purposes. The world doesn't seem to understand that we don't actually have production software infrastructure to scale and manage AI systems to the enormous computational heights required for them to practically work in products across hardware. At Modular, we have talked about these challenges many times, because AI isn't an isolated problem - it's a systems problem comprised of hardware + software across both data centers and the edge.

Regulation: where to start?

So, if we assume that AGI isn't arriving for the foreseeable future, as I and many do, what exactly are we seeking to regulate today? Is it the models we have already been using for many years in plain sight, or are we trying to pre-empt a future far off? As always, the truth lies somewhere in between. It is the generative AI revolution - not forms of AI that have existed for years - that has catalyzed much of the recent excitement around AI. It is here, in this subset of AI, where practical AI regulation should take a focused start.

If our guiding success criteria is broadly something like - enable AI to augment and improve human life - then let's work backward with clarity from that goal and implement legislative frameworks accordingly. If we don't know what goal we are aiming for - how can we possibly define policies to guide us? Our goal cannot be "regulate AI to make it safe" because the very nature of that statement infers the absence of safety from the outset despite us living with AI systems for 10+ years. To realistically achieve a balanced regulatory approach, we should start within the confinements of laws that already exist - enabling them to address concerns about AI and learn how society reacts accordingly - before seeking completely new, sweeping approaches. The Venn diagram for AI and existing laws is already very dense. We already have data privacy and security laws, discrimination laws, export control laws, communication and content usage laws, and copyright and intellectual property laws, among so many other statutory and regulatory frameworks today that we could seek to amend.

The idea that we need entirely new "AI-specific" laws now - when this field has already existed for years and risks immediately curbing innovation for use cases we don't yet fully understand - feels impractical. It will likely be cumbersome and slow to enforce while creating undue complexity that will likely stifle rather than enable innovation. We can look to history for precedent here - there is no single "Internet Act" for the US or the world - instead, we divide and conquer Internet regulation across our existing legislative and regulatory mechanisms. We have seen countless laws that attempt to broadly regulate the internet fail - the Stop Online Piracy Act (SOPA) is one shining US example - while other laws that seek to regulate within existing bounds succeed. For example, Section 230 of the Communications Decency Act protects internet service providers from being liable for content published, and the protection afforded here has enabled modern internet services to innovate and thrive to enormous success (e.g., YouTube, TikTok, Instagram etc.) while also forcing market competition on corporations to create their own high content standards to build better product experiences and retain users. If they didn't implement self-enforcing policies and standards, users would simply move to a better and more balanced service, or a new one would be created - that's market dynamics.

Of course, we should be practical and realistic. Any laws we amend or implement will have failings - they will be tested in our judicial system and rightly criticized, but we will learn and iterate. Yes, we won’t get this right initially - a balanced approach often will mean some bad actors will succeed, but we can limit these bad actors while dually enabling us to build a stronger and more balanced AI foundation for the years ahead. AI will continue evolving, and the laws will not keep pace - take, for example, misinformation, where Generative AI makes it far easier to construct alternate truths. While this capability has been unlocked, social media platforms still grapple with moderation of non-AI-generated content and have done so for years. Generative AI will likely create extremely concerning misinformation across services, irrespective of any laws we implement.

EU: A concerning approach

With this context, let’s examine one of the most concerning approaches to AI regulation - the European Artificial Intelligence Act - an incredibly aggressive approach that will likely cause Europe to fall far behind the rest of the world in AI. In seeking to protect EU citizens absolutely, the AI Act seemingly forgets that our world is now far more interconnected and that AI programs are, and will continue to be, deeply proliferated across our global ecosystem. For example, the Act arguably could capture essentially all probabilistic methods of predicting anything. And, in one example, it goes on to explain (in Article 5) that:

(a) the placing on the market, putting into service or use of an AI system that deploys subliminal techniques beyond a person’s consciousness in order to materially distort a person’s behaviour in a manner that causes or is likely to cause that person or another person physical or psychological harm;

Rather than taking a more balanced approach and appreciating that perfect is the enemy of good, the AI Act tries to ensnare all types of AI, not just the generative ones currently at the top of politicians and business leaders' minds. How does one even hope to enforce "subliminal techniques beyond a person's consciousness" - does this include the haptic notification trigger on my Apple Watch powered by a probabilistic model? Or the AI system that powers directions on my favorite mapping service with different visual cues? Further, the Act also includes a "high risk" categorization for "recommender systems" that, today, basically power all of e-commerce, social media, and modern technology services - are we to govern and require conformity, transparency and monitoring of all of these too? The thought is absurd, and even if one disagrees, the hurdle posed for generative AI models is so immense - no model meets the standards of the EU AI Act in its current incarnation.

We should not fear AI - we've been living with it for 10+ years in every part of our daily lives - in our news recommendations, mobile phone cameras, search results, car navigation, and more. We can't just seek to extinguish it from existence retrospectively - it's already here. So, to understand what might work, let's walk the path of history to explore what didn't - the failed technological regulations of the past. Take the Digital Millennium Copyright Act (DMCA), which attempted to enforce DRM by making it illegal to circumvent technological measures used to protect copyrighted works. This broadly failed everywhere, didn't protect privacy, stifled innovation, was hacked, and, most critically - was not aligned with consumer interests. It failed in Major League Baseball, the E-Book industry abandoned it, and even the EU couldn't get it to pass. What did work in the end? Building products aligned with consumer interests enabled them to achieve their goal - legitimate ways to access high-quality content. The result? We have incredible services like Netflix, Spotify, YouTube, and more - highly consumer-aligned products that deliver incredible economic value and entertainment for society. While each of these services has its challenges, at a broad level, they have significantly improved how consumers access content, have decentralized its creation, and enabled enormous and rapid distribution that empowers consumers to vote on where they direct their attention and purchasing power.

A great opportunity to lead

The US has an opportunity to lead the world by constructing regulation that enables progressive and rapid innovation. It should be merited by this country's principles and its role as a global model for democracy. Spending years constructing the "perfect AI legislation" and "completely new AI agencies" will end up like the parable of the blind men and an elephant - creating repackaged laws that attempt to regulate AI from different angles without holistically solving anything.

Ultimately, seeking to implement broad new statutory protections and government agencies on AI is the wrong approach. It will be slow, it will take years to garner bipartisan support, and meanwhile, the AI revolution will roll onwards. We need AI regulation that defaults to action, that balances opening the door to innovation and creativity while casting a shadow over misuse of data and punishing discriminative, abusive, and prejudicial conduct. We should seek regulation that is more focused initially, predominantly on misinformation and misrepresentation, and seek to avoid casting an incredibly wide net across all AI innovation so we can understand the implications of initial regulatory enforcement and how to structure it appropriately. For example, we can regulate the data on which these models are trained (e.g. privacy, copyright, and intellectual property laws), we can regulate computational resources (e.g. export laws), and we can regulate what products predictions are made in (e.g. discrimination laws) to start. In this spirit, we should seek to construct laws that target the inputs and outputs of AI systems and not individual developers or researchers who push the world forward.

Here’s a small non-exhaustive list of near-term actionable ideas:

  1. Voluntary transparency on the research & development of AI - If we wait for Congress, we wait for an unachievable better path at the expense of a  good one today. There is already an incredible body of work being proposed by Open AI, in Google's AI Principles among others on open and transparent disclosure - there is a strong will to transcend Washington and just do something now. We can seek to ensure companies exercise a Duty of Care, to have a responsibility under common law to identify and mitigate ill effects and have transparent reporting accordingly. And of course, this should extend and be true of our Government agencies as well.

  2. Better AI use case categorization & risk definition - There is no well-defined regulatory concept of “AI” nor how to “determine risks” on “what use cases” therein. While the European AI Act clearly goes too far, it does at least seek to define what “AI is” but errs in defining essentially all of probability. Further, it needs to better classify risk and the classes of AI from which it is actually seeking to protect people. We can create a categorical taxonomy of AI use cases and seek to target regulatory enforcement accordingly. For example, New York is already implementing laws to require employers to notify job applications if AI is used to review their applications.

  3. Pursue watermarking standards - Google leads the cause with watermarking to determine whether content has been AI generated. While these standards are reminiscent of DRM, it is a useful step in encouraging open standards for major distribution platforms to enable watermarking to be baked in.

  4. Prompt clearly on AI systems that collect and use our data - Apple did this to great effect, implementing IDFA and better App Transparency on how consumer data is used. We should expect more of our services for any data being used to train AI along with the AI model use cases such data informs. Increasingly in AI, all data is incredibly biometric - from the way one types, to the way one speaks, to the way one looks and even walks. All of this data can, and is, now being used to train AI models to form biometric fingerprints of individuals - this should be made clear and transparent to society at large. Both data privacy and intellectual property laws should protect who you are, and how your data is used in a new generative AI world. Data privacy isn’t new, but we should realize that the evolution of AI continually raises the stakes.

  5. Protect AI developers & open source - We should seek to focus regulatory efforts on the data inputs as well as have clear licensing structures for AI Models and software. Researchers and developers should not be held liable for creating models, software tooling, and infrastructure that is distributed to the world so that an open research ecosystem can continue to flourish. We must promote initiatives like model cards and ML Metadata across the ecosystem and encourage their use. Further, if developers open source AI models, we must ensure that it is incredibly difficult to hold the AI model developer liable, even as a proximate cause. For example, if an entity uses an open-source AI recommendation model in its system, and one of its users causes harm as a result of one of those recommendations, then one can’t merely seek to hold the AI model developer liable - it is foreseeable to a reasonable person who develops and deploys AI systems that rigorous research, testing, and safeguarding must occur before using and deploying AI models into production. We should ensure this is true even if the model author knew the model was defective. Why, you ask? Again, a person who uses any open source code should reasonably expect that it might have defects - hence why we have licenses like MIT that provide software “as is”. Unsurprisingly, this isn’t a new construct; it's how liability has existed for centuries in tort, and we should seek to ensure AI isn't treated differently.

  6. Empower agencies to have agile oversight - AI consumer scams aren’t new - they are already executed by email and phone today. Enabling existing regulatory bodies to have the power to take action immediately is better than waiting for congressional oversight. The FTC has published warnings, the Department of Justice as well, and the Equal Employment Opportunity Commission too - all of the regulatory agencies that can help more today.

Embrace our future

It is our choice to embrace fear or excitement in this AI era. AI shouldn’t be seen as a replacement for human intelligence, but rather a way to augment human life - a way to improve our world, to leave it better than we found it and to infuse a great hope for the future. The threat that we perceive - that AI calls into question what it means to be human - is actually its greatest promise. Never before have we had such a character foil to ourselves, and with it, a way to significantly improve how we evolve as a species - to help us make sense of the world and the universe we live in and to bring about an incredibly positive impact in the short-term, not in a long distant theoretical future. We should embrace this future wholeheartedly, but do so with care and great tact - for as with all things, life is truly a great balancing act.

Many thanks to Christopher Kauffman, Eric Johnson, Chris Lattner and others for reviewing this post.

Image credits: Tim Davis x DALL-E (OpenAI)

Read More
Interviews Tim Davis Interviews Tim Davis

The Aussie conquering Silicon Valley

Meet the little-known young Australian in Silicon Valley at the forefront of the artificial intelligence revolution, who heads one of the hottest start-ups in America that has the audacious goal of fixing AI infrastructure for the world’s software and hardware developers.

By JOHN STENSHOLT (The Australian Business Review)

After years of toughing it out and watching his Silicon Valley dream almost die, a little-known former Melbourne NAB analyst now leads one of the hottest start-ups in America.

Meet the little-known young Australian in Silicon Valley at the forefront of the artificial intelligence revolution, who heads one of the hottest start-ups in America that has the audacious goal of fixing AI infrastructure for the world’s software and hardware developers.

Tim Davis, 40, turned up in California on a whim a little over a decade ago, survived a stint in a wild house known as the “Hacker Fortress”, started an early version of an online food delivery service before the rise of UberEats and others, was poached by Google only to leave the technology giant in 2022 to co-found Modular, an AI infrastructure start-up recently valued at about US$600 million (A$927 million).

And he says he is only getting started.

In his first Australian media interview, Davis says the goals for Modular are clear – and big.
“We started Modular to improve AI infrastructure for the world. Changing the world is never easy – but we are incredibly determined to do so,” he said.

“AI is so important to the future of humanity, and we feel a great purpose to try to improve AI’s usability, scalability, portability and accessibility for developers and enterprises around the world.”

Modular’s vision is to allow AI technology to be used by anyone, anywhere and it is creating a developer infrastructure platform that enables more developers around the world to deploy AI faster, and across more hardware.

Davis, and his America co-founder Chris Lattner, have helped build and scale much of the AI software infrastructure that powers workloads at some of the world’s largest tech companies – including Google – but they argue this software has many shortcomings and was designed for research and not for scaling AI across the vast number of new uses and hardware the world is demanding now and into the future.

They are aiming to rebuild and unify AI software infrastructure, solving fragmentation issues that make it difficult for developers who work outside the world’s largest companies to build, deploy and scale AI, and ultimately make AI more accessible to everyone.

“We thought we could build something unique that could actually empower the world to move faster with AI, while equally making it more accessible to developers, make it easier to program in and better from a cost standpoint because you can scale it to different types of hardware.”

Modular claims to have built the world’s fastest AI inference engine – software that enables AI programs to run and scale to millions of people – and its own programming language, Mojo, a superset of Python (the world’s most popular programming language) which enables developers to deploy their AI programs tens of thousands of times faster, reduce costs and make it more simple to deploy AI around the world.

After raising US$30 million from investors last year, Modular recently raised another US$100 million in a funding round led by private equity firm General Catalyst and including Google Ventures, and Silicon Valley-based venture funds SV Angel, Greylock Partners and Factory.

Modular says it’s now more than 120,000 developers using its products—such as its inference engine and Mojo—launched in early May—and that “leading tech companies” are already using its infrastructure, with 35,000 on the waitlist.

The US$100 million raised will be used on product expansion, hardware support and the expansion of Mojo, as well as building the sales and commercialisation aspects of the business.

Early years

All of which is a far cry from when Davis flew to Silicon Valley in September 2012 with dreams of becoming an entrepreneur, following stints working as a financial analyst at National Australia Bank and then several internships at law firms like Allens and Minters after completing law, business and commerce degrees at Melbourne’s Monash University.

Davis had started his own company called CrowdSend in 2011, which had a software system for identifying objects in images and pictures and matching them to retailers, but said he found it “difficult” as “investors in Australia, if you wanted to raise some seed capital, they would take a very large percentage of the business.”

“So it was my wife that said if you want to do this (become an entrepreneur) why don’t you go do the place where technology is. We’d actually just gotten married and then off I went to America.”

Davis used Airbnb to find a place to stay in Silicon Valley, gave himself six weeks to make a success of things and found what was called the Hacker Fortress in the Los Altos Hills.

“It was very much a place with a whole bunch of misfits who had come to America, particularly Silicon Valley, to do a start-up. This was a very big house, it had 15 rooms, and my dream at the time was how to make CrowdSend successful. There were a lot of really talented people there,” he says.

“(But) it turns out when you come to Silicon Valley, and you come with a preconceived notion of what you want to do, well then there’s reality of what you ended up doing.”

There were plenty of issues in the Hacker Fortress, which was not as clean as advertised on Airbnb. At one stage, a housemate died in his room. Davis would also have the roof fall in on his room one evening.

The house, despite being advertised as being only 10 minutes from the likes of Google and Apple, was actually quite a way from shops and restaurants.

“We basically had no ability to get food easily and so we had this idea that maybe we could build this large scalable distribution model,” Davis explains. “At the time the likes of GrubHub and these other businesses were going to restaurants and trying to get commission deals. So instead we reverse engineered the Starbucks and Chipotle menus and built an app and threw it out as an idea to people in the house.”

Before UberEats

The idea for what would become Davis’s next business, Fluc (Food Lovers United Co), was born – all because he and his housemates were a little lazy and really didn’t want to cook.

It was 2013, a year before UberEats was launched, and while other food delivery service apps were dealing directly with restaurants, Davis and his took a different approach.

“We just thought, why don’t we just grab every restaurant menu, we just put it on our website, inflate the prices, and start selling food? Overnight, we went from five restaurants to 160 restaurants. And it just exploded. It was unprecedented. And what we tapped into really was that selection mattered. And that’s what consumers had not had, at least in the American market. And so then we went on this journey of raising capital.”

Fluc would mainly service the local Bay Area market of about 8 million people, though it tried to move to Los Angeles at one stage, and raised US$4 million from local angel investors.

Davis says after two and a half years of working seven days a week it became obvious “that we wouldn’t be able to raise the amount of capital that we needed to scale that business” at a time when the likes of DoorDash were raising hundreds of millions of dollars annually from investors.

Davis and his team then started talking to other companies about potentially being acquired, and eventually Google expressed interest not in the company but the people Fluc had assembled.

That led to Davis joining Google, where he worked as a product manager and then joined the ad division where he built machine learning systems and the Google Brain research division dedicated to AI – where he met Lattner.

What followed, Davis says, was six years of building AI systems that stood them in good stead when they decided to leave Google and form Modular.

“We were basically there for all the major components of AI as it is now … and what was interesting is that through that experience we could see the infrastructure was designed by researchers – a bunch of people who wanted to train large machine learning models. But taking those models from a research environment and actually scaling them into very large production systems is very hard.”

He says Modular’s only goal is to find solutions to that problem, and commercialise it and that the opportunity “we have in this market is astronomically huge.”

What’s next

“We work with everyone from high performance racing teams, to autonomous car companies, to very large machine learning recommendations to generative AI. You name it, whether it’s video generation, image generation, text generation, we’re there.

“You could go down the NASDAQ list of companies (to find those) who want to use our infrastructure to scale AI inside their organisations.”

As for widespread concerns about AI, Davis says the fact that AI systems take an “incredible amount of time to build” means that the public should be “realistic” about the perception that AI could be a threat to humanity.

“We definitely have better recommendation systems, we have better chat bots, we have translation across multiple languages, we can take better photos, we are now all better copywriters, and we have voice assistants.

“But none of this is remotely close to AGI or artificial general intelligence – or anywhere near ASI: artificial super intelligence. We have a long way to go.

“AI shouldn’t be seen as a replacement for human intelligence, it’s a way to augment human life – a way to improve our world, to leave it better than we found it.”

Read More
Interviews Tim Davis Interviews Tim Davis

Unite AI – Interview Series

Tim Davis, is the Co-Founder & President of Modular, an integrated, composable suite of tools that simplifies your AI infrastructure so your team can develop, deploy, and innovate faster. Modular is best known for developing Mojo, a new programming language that bridges the gap between research and production by combining the best of Python with systems and metaprogramming.

By Antoine Tardif (Unite AI)

Tim Davis, is the Co-Founder & President of Modular, an integrated, composable suite of tools that simplifies your AI infrastructure so your team can develop, deploy, and innovate faster. Modular is best known for developing Mojo, a new programming language that bridges the gap between research and production by combining the best of Python with systems and metaprogramming.

Repeat Entrepreneur and Product Leader. Tim helped build, found and scale large parts of Google's AI infrastructure at Google Brain and Core Systems from APIs (TensorFlow), Compilers (XLA & MLIR) and runtimes for server (CPU/GPU/TPU) and TF Lite (Mobile/Micro/Web), Android ML & NNAPI, large model infrastructure & OSS for billions of users and devices. Loves running, building and scaling products to help people, and the world.

When did you initially discover coding, and what attracted you to it?

As a kid growing up in Australia, my dad brought home a Commodore 64C and gaming was what got me hooked – Boulder Dash, Maniac Mansion, Double Dragon – what a time to be alive. That computer introduced me to BASIC and hacking around with that was my first real introduction to programming. Things got more intense through High School and University where I used more traditional static languages for engineering courses, and over time I even dabbled all the way up to Javascript and VBA, before settling on Python for the vast majority of programming as the language of data science and AI. I wrote a bunch of code in my earlier startups but these days, of course, I utilize Mojo and the toolchain we have created around it.

For over 5 years you worked at Google as Senior Product Manager and Group Product Leader, where you helped to scale large parts of Google's AI infrastructure at Google Brain. What did you learn from this experience?

People are what build world-changing technologies and products, and it is a devoted group of people bound by a larger vision that brings them to the world. Google is an incredible company, with amazing people, and I was fortunate to meet and work with many of the brightest minds in AI years ago when I moved to join the Brain team. The greatest lessons I learnt were to always focus on the user and progressively disclose complexity, to empower users to tell their unique stories to the world like fixing the Greater Barrier Reef or helping people like Jason the Drummer, and to attract and assemble a diverse mix of people to drive towards a common goal. In a massive company of very smart and talented people, this is much harder than you can imagine. Reflecting on my time there, it’s always the people you worked with that are truly memorable. I will always look back fondly and appreciate that many people took risks on me, and I’m enormously thankful they did, as many of those risks encouraged me to be a better leader and person, to dive deep and truly understand AI systems. It truly made me realize the profound power AI has to impact the world, and this was the very reason I had the inspiration and courage to leave and co-found Modular.

Can you share the genesis story behind Modular?

Chris and I met at Google and shipped many influential technologies that have significantly impacted the world of AI today. However, we felt AI was being held back by overly complex and fragmented infrastructure that we witnessed first hand deploying large workloads to billions of users. We were motivated by a desire to accelerate the impact of AI on the world by lifting the industry towards production-quality AI software so we, as a global society, can have a greater impact on how we live. One can’t help but wonder how many problems AI can help solve, how many illnesses cured, how much more productive we can become as a species, to further our existence for future generations, by increasing the penetration of this incredible technology.

Having worked together for years on large scale critical AI infrastructure – we saw the enormous developer pain first hand – “why can’t things just work”? For the world to adopt and discover the enormous transformative nature of AI, we need software and developer infrastructure that scales from research to production, and is highly accessible. This will enable us to unlock the next way of scientific discoveries – of which AI will be critical – and is a grand engineering challenge. With this motivating background, we developed an intrinsic belief that we could set out to build a new approach for AI infrastructure, and empower developers everywhere to use AI to help make the world a better place. We are also very fortunate to have many people join us on this journey, and we have the world's best AI infrastructure team as a result.

Can you discuss how the Mojo programming language was initially built for your own team?

Modular’s vision is to enable AI to be used by anyone, anywhere. Everything we do at Modular is focused on that goal, and we walk backwards from that in the way we build out our products and our technology. In this light, our own developer velocity is what matters to us firstly, and having built so much of the existing AI infrastructure for the world – we needed to carefully consider what would enable our team to move faster. We have lived through the two-world language problem in AI – where researchers live in Python, and production and hardware engineers live in C++ – and we had no choice but to either barrel down that road, or rethink the approach entirely. We chose the latter. There was a clear need to solve this problem, but many different ways to solve it – we approached it with our strong belief of meeting the ecosystem where it is today, and enabling a simpler lift into the future. Our team bears the scars of software migration at large scale, and we didn’t want a repeat of that. We also realized that there is no language today, in our opinion, that can solve all the challenges we are attempting to solve for AI and so we undertook a first principles approach, and Mojo was born.

How does Mojo enable seamless scaling and portability across many types of hardware?

Chris, myself and our team at Google (many at Modular) helped bring MLIR into the world years ago – with the goal to help the global community solve real challenges by enabling AI models to be consistently represented and executed on any type of hardware. MLIR is a new type of open-source compiler infrastructure that has been adopted at scale, and is rapidly becoming the new standard for building compilers through LLVM. Given our team's history in creating this infrastructure, it's natural that we utilize it heavily at Modular and this underpins our state of the art approach in developing new AI infrastructure for the world. Critically, while MLIR is now being fast adopted, Mojo is the first language that really takes the power of MLIR and exposes it to developers in a unique and accessible way. This means it scales from Python developers who are writing applications, to Performance engineers who are deploying high performance code, to hardware engineers who are writing very low level system code for their unique hardware.

References to Mojo claim that it’s basically Python++, with the accessibility of Python and the high performance of C. Is this a gross oversimplification? How would you describe it?

Mojo should feel very familiar to any Python programmer, as it shares Python’s syntax. But there are a few important differences that you’ll see as one ports a simple Python program to Mojo, including that it will just work out of the box. One of our core goals for Mojo is to provide a superset of Python – that is, to make Mojo compatible with existing Python programs – and to embrace the CPython implementation for long-tail ecosystem support. Then enable you to slowly augment your code and replace non-performing parts with Mojo’s lower-level features to explicitly manage memory, add types, utilize autotuning and many other aspects to get the performance of C or better! We feel Mojo gives you get the best of both worlds and you don’t have to write, and rewrite, your algorithms in multiple languages. We appreciate Python++ is an enormous goal, and will be a multi-year endeavor, but we are committed to making it reality and enabling our legendary community of more than 140K+ developers to help us build the future together.

In a recent keynote it was showcased that Mojo is 35,000x faster than Python, how was this speed calculated?

It’s actually 68,000x now! But let's recognize that it's just a single program in Mandelbrot – you can go and read a series of three blog posts on how we achieved this – here, here and here. Of course, we’ve been doing this a long time and we know that performance games aren’t what drive language adoption (despite them being fun!) – it’s developer velocity, language usability, high quality toolchains & documentation, and a community utilizing the infrastructure to invent and build in ways we can’t even imagine. We are tool builders, and our goal is to empower the world to use our tools, to create amazing products and solve important problems. If we focus on our larger goal, it's actually to create a language that meets you where you are today and then lifts you easily to a better world. Mojo enables you to have a highly performant, usable, statically typed and portable language that seamlessly integrates with your existing Python code – giving you the best of both worlds. It enables you to realize the true power of the hardware with multithreading and parallelization in ways that raw Python today can not – unlocking the global developer community to have a single language that scales from top to bottom.

Mojo’s magic is its ability to unify programming languages with one set of tools, why Is this so important?

Languages always succeed by the power of their ecosystems and the communities that form around them. We’ve been working with open source communities for a long time, and we are incredibly thoughtful towards engaging in the right way and ensuring that we do right by the community. We’re working incredibly hard to ship our infrastructure, but need time to scale out our team – so we won’t have all the answers immediately, but we’ll get there. Stepping back, our goal is to lift the Python ecosystem by embracing the whole existing ecosystem, and we aren’t seeking to fracture it like so many other projects. Interoperability just makes it easier for the community to try our infrastructure, without having to rewrite all their code, and that matters a lot for AI.

Also, we have learnt so much from the development of AI infrastructure and tools over the last ten years. The existing monolithic systems are not easily extensible or generalizable outside of their initial domain target and the consequence is a hugely fragmented AI deployment industry with dozens of toolchains that carry different tradeoffs and limitations. These design patterns have slowed the pace of innovation by being less usable, less portable, and harder to scale.

The next-generation AI system needs to be production-quality and meet developers where they are. It must not require an expensive rewrite, re-architecting, or re-basing of user code. It must be natively multi-framework, multi-cloud, and multi-hardware. It needs to combine the best performance and efficiency with the best usability. This is the only way to reduce fragmentation and unlock the next generation of hardware, data, and algorithmic innovations.

Modular recently announced raising $100 million in new funding, led by General Catalyst and filled by existing investors GV (Google Ventures), SV Angel, Greylock, and Factory. What should we expect next?

This new capital will primarily be used to grow our team, hiring the best people in AI infrastructure, and continuing to meet the enormous commercial demand that we are seeing for our platform. Modverse, our community of well over 130K+ developers and 10K’s of enterprises, are all seeking our infrastructure – so we want to make sure we keep scaling and working hard to develop it for them, and deliver it to them. We hold ourselves to an incredibly high standard, and the products we ship are a reflection of who we are as a team, and who we become as a company. If you know anyone who is driven, who loves the boundary of software and hardware, and who wants to help see AI penetrate the world in a meaningful and positive way – send them our way.

What is your vision for the future of programming?

Programming should be a skill that everyone in society can develop and utilize. For many, the “idea” of programming instantly conjures a picture of a developer writing out complex low level code that requires heavy math and logic – but it doesn’t have to be perceived that way. Technology has always been a great productivity enabler for society, and by making programming more accessible and usable, we can empower more people to embrace it. Empowering people to automate repetitive processes and make their lives simpler is a powerful way to give people more time back.

And in Python, we already have a wonderful language that has stood the test of time – it's the world's most popular language, with an incredible community – but it also has limitations. I believe we have a huge opportunity to make it even more powerful, and to encourage more of the world to embrace its beauty and simplicity. As I said earlier, it's about building products that have progressive disclosure of complexity – enabling high level abstractions, but scaling to incredibly low level ones as well. We are already witnessing a significant leap with AI models enabling progressive text-to-code translations – and these will only become more personalized over time – but behind this magical innovation is still a developer authoring and deploying code to power it. We’ve written about this in the past – AI will continue to unlock creativity and productivity across many programming languages, but I also believe Mojo will open the ecosystem aperture even further, empowering more accessibility, scalability and hardware portability to many more developers across the world.

To finish, AI will penetrate our lives in untold ways, and it will exist everywhere – so I hope Mojo catalyzes developers to go and solve the most important problems for humanity faster – no matter where they live in our world. I think that’s a future worth fighting for.

Read More
Short posts Tim Davis Short posts Tim Davis

Modular has raised $100M to fix AI infrastructure

We are so excited to announce this $100M raise, and beyond proud of what our world-class team, incredible customers and partners have enabled us to achieve. Our AI Engine is the world's fastest and has a strong list of customers lining up for its unparalleled performance and usability, and our new programming language in - Mojo 🔥 - already has a developer community of >120K+ developers in just 4 months.

Chris Lattner and I started Modular to help improve AI infrastructure for the world, and to enable the next wave of AI innovation to be truly unlocked on the world's hardware.

We are so excited to announce this $100M raise, and beyond proud of what our world-class team, incredible customers and partners have enabled us to achieve. Our AI Engine is the world's fastest and has a strong list of customers lining up for its unparalleled performance and usability, and our new programming language in - Mojo 🔥 - already has a developer community of >120K+ developers in just 4 months.

This round was led by General Catalyst, in Deep Nishar&Christopher Kauffman, and filled out by existing investors in GV (Google Ventures) with Dave Munichiello, SVA in Steven Lee&Ronny Conway, Greylock in Saam Motamedi and Factory - amazing leaders and incredible people.

We are so fortunate to work with them, and their belief, to build and create change in the world.And of course, changing the world is never easy - but we are so incredibly determined to continue to do so. AI is so important to the future of humanity, and we feel a great purpose to truly improve AI's usability, scalability, portability and accessibility for the worlds developers and enterprises.

Join us on this incredible journey, and let's change the world together 🚀! You can read more on the Modular blog.

Read More
Interviews Tim Davis Interviews Tim Davis

Data Exchange Interview

Welcome to the Data Exchange Podcast. Today we’re joined by Tim Davis, co-founder and Chief Product Officer at Modular. Their tagline says it all: The future of AI development starts here. Tim, great to have you on the show.

Interview with Tim Davis, Co-Founder of Modular. Full interview here

Ben (Host): Welcome to the Data Exchange Podcast. Today we’re joined by Tim Davis, co-founder and Chief Product Officer at Modular. Their tagline says it all: The future of AI development starts here. Tim, great to have you on the show.
Tim Davis: Great to be here, Ben—thanks for having me.

Introducing Mojo: Python, Reimagined

Ben: Let’s dive right in. What is Mojo, and what can developers use today?
Tim: Mojo is a new programming language—a superset of Python, or “Python++,” if you will. Right now, anyone can sign up at modular.com/mojo to access our cloud-hosted notebook environment, play with the language, and run unmodified Python code alongside Mojo’s advanced features.

“All your Python code will execute out of the box—you can then take performance-critical parts and rewrite them in Mojo to unlock 5–10× speedups.”

That uplift comes from our state-of-the-art compiler and runtime stack, built on MLIR and LLVM foundations.

Solving the Two-Language Problem

Many ML frameworks hide C++/CUDA complexity behind Python APIs, but that split still causes friction. Mojo bridges the gap:

  • Prototype in Python

  • Optimize in Mojo (same codebase)

“Researchers no longer need to drop into C++ for speed; they stay in one language from research to production.”

This unified model dramatically accelerates the path from idea to deployment.

Who is Mojo For?

Ben: Frameworks like TensorFlow and PyTorch already tackle performance. Who’s Mojo’s target audience?
Tim: Initially, it’s us—Modular’s own infrastructure team. But our real audience spans:

  1. Systems-level ML engineers who need granular control and performance.

  2. GPU researchers wanting a seamless path to production without rewriting code.

By meeting developers where they are, Mojo helps defragment fragmented ML stacks and simplifies pipelines.

Under the Hood: Hardware-Agnostic Design

Mojo’s architecture is built for broad hardware support:

  • MLIR (Multi-Level IR): Provides a common representation across hardware.

  • LLVM Optimizations: Powers high-performance codegen.

  • Multi-Hardware Portability: CPUs, GPUs, TPUs, edge devices, and beyond.

“We want access to all hardware types. Today’s programming model is constrained—Mojo opens up choice.”

This means you’re not locked into CUDA or any single accelerator vendor.

Beyond the Language: Unified AI Inference Engine

Modular also offers a drop-in inference engine:

  • Integrates with Triton, TF-Serving, TorchServe

  • CPUs first (batch workloads), GPUs coming soon

  • Orders-of-magnitude performance gains

“Simply swap your backend and get massive efficiency improvements—no changes to your serving layer.”

Enterprises benefit from predictable scaling and hardware flexibility, whether on Intel, AMD, ARM-based servers, or custom ASICs.

Roadmap: Community, Open Source & Enterprise

Next 6–12 Months:

  • Expand Mojo’s language features (classes, ownership, lifetimes).

  • Enable GPU execution (beyond the cloud playground).

  • Extend the inference engine to training, dynamic workloads, and full pipeline optimizations (pre-/post-processing).

“We released early to learn from real users—80,000 sign-ups across 230+ countries. Their feedback drives our roadmap.”

Why a New Language Matters

Mojo’s core value prop can be summed up in three words:

  1. Usable: Drop-in Python compatibility; gentle learning curve.

  2. Performant: Advanced compiler + runtime yields 5–10× speedups out of the box.

  3. Portable: Write once, run anywhere—from cloud GPUs to mobile CPUs.

Together, these unlock faster innovation, lower costs, and broader hardware choice.

Democratizing AI Development

In Tim’s own words:

“Our mission is to make AI development accessible to anyone, anywhere. By rethinking the entire stack, we’re unlocking a new wave of innovation and putting compute power in more hands.”

With its unified language and inference engine, Modular is ushering in a future where AI development truly starts here—for researchers, engineers, and enterprises alike.

Read More
Short posts Tim Davis Short posts Tim Davis

Founding Modular & Raising $30M

After working for years in the AI/ML space - I’ve left Google and decided it's time for a new approach to building machine learning infrastructure. Chris Lattner and I, along with an incredible team of talented architects, engineers and product leaders are teaming up to rebuild it from the ground up and truly help the world of AI.

We are bringing together the world's best AI infrastructure talent to improve AI production development and deployment.

After working for years in the AI/ML space - I’ve left Google and decided it's time for a new approach to building machine learning infrastructure. Chris Lattner and I, along with an incredible team of talented architects, engineers and product leaders are teaming up to rebuild it from the ground up and truly help the world of AI.

We building a next generation AI developer platform, and we are proud to partner with @GVteam, @GreylockVC, Factory, SV Angel and notable angels who are funding our $30M first round of funding. We spoke with Dave Munichiello from GV about the opportunity here. The next generation of product breakthroughs will be powered by production quality infrastructure that brings together the best of compilers and runtimes, is designed for heterogeneous compute, edge to datacenter distribution, and is focused on usability. Unifying software and hardware with a "just works" approach that will save developers enormous time and increase their velocity.

Having worked for many years in the AI space at Google, we have and are continuing to assemble the world's best AI infrastructure team. You can read some of the challenges the industry faces, in our opinion, via our blog post here: The Case for a Next-Generation AI Developer Platform. We are hiring for numerous roles - please apply via Modular Careers.

We're excited to showcase what we have been building and designing later in the year. You can checkout a video we put together below.

We are incredibly excited about the mission before us and if you are interested in joining us to change the world - just reach out via www.modular.com. We’re hiring everywhere. The future is super exciting and bright!

Read More