Compute Capital Markets – Multicoin Capital

Bitbuy
Coinmama


Compute is the fastest growing asset class in the world. Over the last few years, there has been an extraordinary amount written about the datacenter buildout, the capital expenditure required to finance token factories of all sizes, and the increasingly creative offtake agreements that AI labs and hyperscalers are signing to secure capacity.

Our first-order conclusion from all of this is that demand for compute will continue to increase, and supply will remain constrained across chips, memory, power and datacenter infrastructure.
The more important implication is that there is now an emergent universe of productive financial assets, which means that there will inevitably be robust markets to price, finance, and hedge exposures around the underlying. There are hundreds of billions of dollars of GPUs sitting inside datacenters, generating cash flows under contracts of different durations and serving workloads with different performance requirements.

The bulk of that flow today is served through bespoke or bilateral agreements between large producers and consumers, and in some cases very disaggregated broker networks or marketplaces. We believe there is an immense opportunity to build the first-order correct exchange infrastructure: a marketplace with embedded financial primitives atop, in order to allow producers and consumers of compute to rigorously express risk preferences along the price and availability curve (and across major categories of hardware). We expect that these new capital markets, which make compute more financialized, are uniquely suited for crypto rails.

In our view, the immediate prize is to build the platform that manages compute exposures across time for the middle to long tail of producers and consumers. Today this happens via a disaggregated network of OTC brokers and desks, but this is the first step on the path to a full fledged exchange. Critically, because of the (largely) non-fungible nature of compute, we believe that all meaningful progress here will begin in mediating the flow through physical delivery, not synthetic or cash-settled instruments. Buying and selling actual units of compute will over time yield some natural clustering or standardization, which will further serve as a base upon which to layer on financial instruments: futures or options contracts, novel capital structures for financing new builds, and ultimately more efficient systems across the lifecycle of a deployment driven by treating the underlying assets as collateral in an open, transparent risk engine.

okex

Crypto rails should be meaningful here for verification, transparency, and composability. A credibly neutral ledger can help turn a verified right to capacity into collateral, stablecoin lenders can lend against that collateral, while a prime broker margins physical inventory and financial positions together. Compute promises to be one of the most exciting domains to extend core financial primitives hardened over DeFi rails into the productive economy.

Why Now?

We believe this is a particularly unique moment to build financial primitives for compute. The first important change is the recent drastic recalibration in residual and resale value for trailing-edge chips driven by inference demand and serving improvements.

For some period, Nvidia’s product cadence led many across the datacenter capital structure to believe that GPUs become economically obsolete within a few years. The argument was something along the lines of: every new architecture represents a step function improvement in performance, therefore older chips will fall out of favor among large consumers of compute, and asset owners depreciating them over five or six years are overstating both earnings and asset values.

We have recently seen strong evidence against this. CoreWeave disclosed that it signed a customer contract for A100 capacity running through 2029. NVidia introduced the A100 in 2020, which means that a customer is willing to commit to using the chip nearly a decade after it was released. CoreWeave also raised a $2.6B credit facility with a five-year maturity against a pool of customer contracts that average three years in duration. The lenders are explicitly underwriting the residual value of the GPUs and CoreWeave’s ability to sell the capacity again after the initial contracts expire.

An A100 may no longer be the chip for training a frontier model, but it can still be useful for inference, fine-tuning, batch processing, academic research, or any of thousands of enterprise workloads for which absolute performance is less important than price or availability. To this effect, you are now seeing substantially more variance in the spot price for A100/H100 chip classes than ever before.

What is becoming abundantly clear is that GPUs will behave less like consumer electronics (where new generations rapidly displace old ones) and more like a ladder of productive assets that serve different segments of demand over time. As chips age, they move down the cost curve and find new buyers, workloads, and geographies.

The second important change is that compute demand is becoming substantially more fragmented.
The first phase of the AI infrastructure buildout was dominated by a small number of frontier labs and hyperscalers. These companies bought massive clusters to train increasingly large proprietary models. They were largely price takers due to the strategic value of cornering supply to compete effectively on the benchmark frontier, but also to monetize later on via inference.

The rise of open-weight models changes this market structure significantly. Today, every enterprise of scale has teams that are working on deploying capable models on their own infrastructure, and inference providers can serve the same model across many different GPU configurations. End customers choose between dozens of models based on cost, latency, geography, and performance. Sovereign AI initiatives add another dimension because governments and regulated enterprises increasingly want models deployed within specific jurisdictions and under specific security or data residency requirements.

The result is a much larger number of buyers with much more heterogeneous demand.
The third change is the emergence of model routing. OpenRouter is an early example of what an order flow auction for inference might look like. Applications submit requests, and the router determines which model and provider should serve them based on price and latency. Underlying providers can charge dramatically different prices. For example, OpenRouter recently showed that Llama 3.3 70B input pricing ranged from $0.10 to more than $1.00 per million tokens across providers.

Arguably the price dispersion reflects that inference providers are less like wholesalers of compute, and more like market makers. They generate margin through batching, caching, quantization, model placement, and most importantly the ability to maintain high utilization across many customers. Two providers with the same hardware can produce a drastically different number of tokens per dollar, in the same way that two refiners can generate different economics from the same barrel of crude oil. These aggregators are an important source of price discovery because they show in real time the marginal value of different models, providers, and hardware configurations.

These developments complicate the compute market, but they also make financial markets around compute considerably more valuable. There are more buyers, more sellers, more configurations, more persistent price dispersion, and a longer economic life over which the underlying assets can be contracted and financed.

Why Don’t Financial Markets For Compute Exist Already?

There has been much discussion recently about whether compute is a commodity or not, and whether traditional financial instruments around commodities can be mapped on to compute parri passu. We don’t believe that this kind of analogy is useful at this stage of the market.

Consider a unit of oil within a particular grade, broadly interchangeable with another unit of the same grade. Compute does not have this property because the value of a GPU depends on the configuration and environment in which it is deployed, and the particular state that device handles.

Take the number of variables embedded in a seemingly simple H100 contract:

  1. PCIe or SXM
  2. The number of GPUs in the node
  3. Available memory
  4. Interconnect, Networking topology
  5. Storage, CPU, geographic location
  6. Security, uptime, support, and the duration and contiguity of the reservation.

An H100 available for a few hours is a different product from an H100 available continuously for three months. A cluster in Virginia is different from one in Iceland, even if the GPUs themselves are identical. Similar hardware can trade for $2 per GPU hour in one market and $15 in another. Some portion of the difference represents market fragmentation, but a larger portion reflects differences in the product itself. Further, reservations are notoriously difficult to unwind because compute is operationally embedded: workloads carry state, data, dependencies, and performance requirements, so moving them between providers requires heavy migration work rather than simply handing another buyer an SSH key.

Today, most buyers manage compute exposure through some combination of spot capacity and long-term reservations. Spot and on-demand contracts preserve flexibility, while forcing buyers to accept uncertainty around price and availability. Long-term reservations guarantee access, while forcing buyers to accept the risk that they overestimate their needs and leave expensive capacity unused.

The problem around compute is less a supply crisis, and more a duration mismatch. This becomes most visible during periods of variable demand. An inference provider may operate comfortably for months and then experience a large demand spike when one of its customers launches a product. A consumer AI application may suddenly go viral and require several times its normal capacity. An enterprise may need a large cluster for a fine-tune or evaluation that lasts only a few weeks. In each case, the buyer faces a choice between overcommitting to long-term capacity or relying on a spot market that may fail at precisely the moment capacity becomes most valuable.

The largest buyers solve this problem through scale. They acquire far more capacity than any single workload requires, maintain inventory across multiple hardware generations and datacenters, and continuously reallocate workloads across their portfolio. Their size allows them to absorb underutilization and manage demand internally.

For the middle and longer tail of the market, access to compute can be existential. These companies cannot afford to carry large quantities of unused capacity, but they also cannot afford to discover that capacity is unavailable when their business needs it. Many of the consumer AI applications of the last several years proved that it is possible to have enormous user demand and still have a structurally broken business because inference costs and capacity requirements were not managed correctly.

Despite this, buyers have almost no ability to hedge either the price or availability of compute. Sellers have limited ways to monetize future capacity, transfer existing obligations, or finance GPUs without producing a long-term customer contract. The market has created the physical assets and the commercial demand, but it has not created standardized claims against those assets.

Compute is inherently physical

The most straightforward attempt to financialize compute, or any asset for that matter, is to construct a price index and list cash-settled futures against it. This gives participants a reference price and allows them to express a view on whether the average cost of a particular GPU will increase or decrease. We have seen early proposals and implementations of these across exchanges like Lighter and Architect.

However, the basis between the reference contract and the physical product is likely to be enormous.

A buyer that needs a B200 cluster in Europe for a continuous 90-day period receives limited protection from a cash-settled future based on the average hourly price of an H100 in the United States. Even within a single GPU generation, differences in networking, memory, location, contract duration, and service levels can cause the physical price paid by the buyer to move independently of the index.

Cash settlement also creates a capital-efficiency problem. Neoclouds already borrow heavily to finance GPUs and datacenter infrastructure, while inference providers need working capital to pay for compute before they collect revenue from customers. Requiring these companies to post additional cash margin against a futures position makes hedging more expensive for the natural participants.

Finally, a cash-settled market without a deep physical market underneath it will primarily attract speculators. Speculators are useful sources of liquidity, but they cannot reliably enforce convergence between the price of a financial contract and the value of actual deliverable capacity. That requires market makers who can procure, deliver, substitute, and transform physical compute.

As a result, we are confident that any set of meaningful financial primitives that serve producers and consumers have to be built atop physical delivery.

Desk to Exchange: Building The Compute Prime Broker

The existing compute market is dominated by hyperscalers, direct sales teams, brokers, and a growing number of spot marketplaces. Brokers facilitate bespoke transactions and rely heavily on relationships. Marketplaces aggregate inventory and expose more prices, but generally do not standardize contracts or become the counterparty to trades.

We believe the path toward compute capital markets begins with a principal desk that sits between buyers and sellers, structures contracts, manages physical delivery, and gradually standardizes the most common forms of capacity. By intermediating real transactions, the desk develops proprietary information about configuration-level pricing, counterparty quality, utilization, delivery failures, and the actual basis between different types of compute.

This information should make it possible to create transferable claims against physical capacity.
For example, a seller could commit one eight-GPU H100 SXM node in US East for a 30-day period and receive a standardized receipt representing that capacity. The receipt could be transferred before the reservation begins, used as collateral, or delivered into a forward contract. At maturity, the holder receives access to the underlying cluster.

It is our assessment that compute will never become completely fungible, but fungibility is not binary. Two clusters can be sufficiently interchangeable for a particular class of workloads even if they are not substitutes for every buyer.

For example, many inference providers serving open source models may accept an eight-GPU H200 SXM node with NVLink from several providers, provided it is available in an acceptable region for a continuous 30 day period and meets some threshold uptime requirements. We expect recurring demand to cluster around a limited number of configurations and reservation windows, giving the market natural starting points for standardization. The principal desk can discover these groupings through actual transactions, and verify that capacity from different sellers meets the same delivery standard. If there are enough buyers that treat any number of clusters as effectively substitute goods, future capacity can trade against some common reference contract around that cluster, thereby creating a Schelling point for liquidity.

Once a venue sits in the physical flow, there is a wide design space for financial products.

  • Transferable forwards: Buyers can lock in future capacity and later sell their position if their needs change. Sellers receive greater revenue visibility without permanently binding capacity to a single customer.
  • Capacity options: Buyers can purchase the right to access a specified cluster during a future period, which is especially valuable for product launches, model releases, fine-tunes, evaluations, and other variable workloads.
  • Portfolio margin: A prime broker can recognize offsetting exposures across spot inventory, forwards, options, hardware generations, locations, and durations. A provider that is long H100 capacity and short B200 capacity should be margined against the risk of the spread rather than the gross value of both positions.
  • Lit RFQs and order books: Standard contracts can trade transparently, while large and bespoke requirements are matched through competitive RFQs. The two structures can coexist because buyers of $100M of annual compute have fundamentally different procurement needs from buyers of $1M.
  • Verification: The venue can verify that capacity exists, meets the represented performance characteristics, and has not been sold to multiple buyers. For inference markets, the same verification layer can measure tokens, latency, throughput, and uptime to prevent providers from spoofing performance.
  • Financing: Smaller neoclouds can borrow against certified inventory, transferable forwards, and hedged future capacity rather than relying exclusively on committed customer revenue. A visible forward curve also gives lenders a better framework for underwriting residual value and renewal risk.

These products collectively resemble a compute prime broker or merchant bank.

AWS and CoreWeave already perform many of these functions internally. They commit to datacenters and chips before selling the capacity across several durations. Their portfolio of assets, and deep demand side liquidity, lets them manage the resulting mismatch.

A smaller neocloud is effectively an undercapitalized merchant. It borrows to purchase GPUs and sells contracts of different durations, absorbs underutilization, and takes on residual value risk. Each company manages these exposures independently even when they could be netted across the market.

An independent prime broker can aggregate risk across the market by warehousing one exposure and netting it against an offsetting position. A diversified pool of these could then theoretically support new forms of credit. This type of market participant seeks to transmute the “shape mismatch” that compute has today, thereby reducing idle capacity and lowering the cost of capital for new financing.

Over time, the reference configurations may actually support credible price indices and useful cash-settled derivatives. Their effectiveness as hedging instruments depends on how closely the index actually tracks some set of repeatable workloads from a large enough set of buyers. Ultimately, even a perfect price hedge leaves the buyer responsible for securing usable capacity when it is needed. Standardization therefore makes both forms of contract more useful: physical delivery solves for availability risk, cash settlement handles price risk.

Compute as Collateral (on Crypto Rails)

The core primitive could look like a credible digital claim against a physical asset or future service.

Examples of this could include:

  1. A GPU or a secured interest in a server
  2. The right to use a cluster during a future window
  3. Receivables under an offtake agreement, pooled into a standardized claim

The purpose of tokenization here would be to make underlying assets and associated cash flows easier to transfer, pledge, and rehypothecate as collateral. In every case, it’s important that there is a strong legal right underneath the claims. This is, in and of itself, an opportunity rich design space for regulatory and verification infrastructure that is a prerequisite to robust capital markets around this asset class.

It’s possible to imagine that, at scale, a neocloud could issue a receipt against verified future capacity and sell part of it forward. The remaining inventory could support a stablecoin denominated margin loan. A lender could underwrite the receipt using its forward price and contracted cash flow, and utilization history would inform the haircut. If the borrower defaults, the collateral can move to a market maker that already knows how to deliver or resell the compute.

A buyer could purchase a call option for a product launch and resell it if plans change, and exercise would settle through delivery versus payment. An inference provider could finance a reservation against customer revenue and hold a capacity option in the same margin account. A prime broker could accept a capacity receipt and extend credit against it, and the receipt could then support a broader secured funding arrangement.

The central actor in the immediate term is a “compute prime broker,” which today is fundamentally constructing its own collateral ledger. It records the asset and the obligation against it, and shows where each claim sits in the capital stack before handling margin and settlement. At scale, we believe that DeFi provides a natural architecture for this market with many asset originators and many sources of capital.

There are a few reasons for this:

  1. Collateral is transparent: A lender can see whether a capacity receipt has already been pledged and which debt is senior. The haircut and maturity are visible on the same ledger.
  2. Collateral is composable: A capacity receipt held in a lending pool can also serve as margin on the forward that protects its value. Both venues rely on the same custody layer.
  3. Settlement is programmable because compute is delivered over time: Payment can be released as capacity becomes available. Collateral returns once the service level is met; failure triggers the agreed penalty. With the physical facts attested, money and collateral can move automatically.
  4. Crypto rails also reach global capital that lends against digital collateral around the clock: Stablecoins provide a common settlement asset (e.g., one vault can hold senior credit while another finances inventory or options).

Credibly-neutral blockchains can coordinate the claim and its settlement, and cryptographic hardware attestations can verify what exists. Each of these, layered over legally-necessary agreements, establish ownership and enforcement, and can form the substrate necessary to unlock capital efficiency across the space.

Compute already supports enormous assets and credit, with very little infrastructure for transferring risk. Ultimately, we believe that protocols don’t capture value, DAOs manage risk. The protocols that build the intermediary between producers and consumers can turn those contracts into a transparent collateral system.

We think this is an incredibly rich design space, and one that is evolving extremely fast. Please reach out to shayon@multicoin.capital if you are working on these problems or have a different perspective on how this market should develop.

Thanks to Kyle Morris for the many discussions that helped shape the thinking behind this post, and to Tomasz Tunguz and Mason Nystrom for their feedback.



Source link

Blockonomics

Be the first to comment

Leave a Reply

Your email address will not be published.


*