Steve Hou GPU Rental Rates, Token Price Trends, AMD’s Inference Push, and the AI Bubble.
Subscribe: Apple • Spotify • Overcast • Pocket Casts • YouTube • AntennaPod • Podcast Addict • Amazon • RSS.
Steve Hou, Head of Research at Silicon Data, joins Ben Lorica to unpack the data behind the AI compute market: GPU rental indices, forward curves, and token expenditure trends that reveal a market still tightening despite talk of oversupply. They cover AMD and Cerebras’s inference ambitions, why open-weight models are pulling down effective token costs even as list prices hold steady, and the financial risks lurking beneath AI’s debt-fueled data center buildout.
Interview highlights – key sections from the video version:
-
-
- Introduction to Silicon Data and Its AI Compute Benchmarks
- Bringing Transparency and Price Discovery to AI Compute
- How Silicon Data Builds and Normalizes GPU Price Indices
- AMD, Cerebras, and TPUs in the Inference Hardware Race
- Token Prices, Task Costs, and the Shift Toward Cheaper Models
- What Market Data Can and Cannot Reveal About Model Usage
- Post-Training, Private Models, and Hidden Inference Demand
- NVIDIA GPU Rentals, Forward Curves, and Rising Residual Values
- AI Data-Center Economics, Inference Margins, and Debt Financing
- Specialized Models and the Rise of Multimodel Enterprise AI
- China’s Emerging AI Hardware Ecosystem and Enterprise Demand
- Regional Compute Demand and the Global GPU Market
- The RAM Boom and Whether High Margins Can Last
- Custom Inference Chips and the Next Wave of AI Hardware
- Power Constraints, Smaller Clusters, and Distributed AI Infrastructure
-
Related content:
- A video version of this conversation is available on our YouTube channel.
- AI Can Win While Data Centers Lose
- Open Models Will Absorb Most of the AI Spend
- Specialized AI Is Getting Easier to Build
- Three New Models, One Signal About Where AI Spending Goes Next
- Manos Koukoumidis: Stop Renting Generic Intelligence for Your Business
- Ameet Talwalkar: Why Observability May Be AI’s Next Frontier
Support our work by subscribing to our newsletter📩
Transcript
Below is a polished and edited transcript.
Ben Lorica. All right, today we’re in for a treat. We’re joined by Steve Hou, head of research at Silicon Data. The website is silicondata.com, and the tagline is “Compute Market Intelligence for Better Decisions: Real-Time Price Transparency and GPU Performance Data for Traders, Financial Institutions, and Builders.” Steve, welcome to the podcast.
Steve Hou. Thank you very much, Ben, for having me.
Ben Lorica. We’re going to have an interesting discussion about the current state of AI. But before we do that, let me try to describe for our audience the types of data you collect. Feel free to correct me at any point, because this will provide context for some of your opinions during the podcast.
One dataset is an independent GPU rental price index: What does it cost to rent a GPU? You also have a RAM index. You have an LLM token expenditure index, which provides a blended rate, in dollars per million tokens, for what it actually costs to run inference. And one of my favorite things is the GPU forward curve, which is like a yield curve for GPUs, although that might be too complex for our audience.
Steve Hou. I love simplifying complex things, so I’m happy to talk about all of them.
Ben Lorica. Generally, I’m trying to paint a picture that you collect data on the cost of hardware and the cost of running inference. Is that a fair description of what you do, or am I missing any big categories?
Steve Hou. That’s right. We’re trying to bring data transparency and price discovery to the physical AI compute market. This is a huge, fast-growing market, but it’s also still very young and opaque.
Carmen, the founder and CEO of Silicon Data, and I both spent time at Bloomberg. We bring that Bloomberg gene with us in this effort, for better or worse.
It’s a complex and fast-evolving industry. There are many different layers and units of account over which people can contract. We focus on the dimensions that most people actually care about. The hourly GPU rental rate is one. You also mentioned the token index. Within GPUs, there are different types of GPUs and different contract terms. Those are the dimensions we focus on.
Ben Lorica. In a previous lifetime, I was a quant at a hedge fund, where I designed trading models for derivatives. I was a Bloomberg Terminal user. I’m assuming you operate somewhat like Bloomberg, at least as it worked back in the day: You’re on the terminal, you notice that a price quote seems wrong, you pick up the phone, someone answers, and they correct it. I’m assuming you’re equally obsessive about data quality and accuracy.
Steve Hou. “Obsessive” might not be the word we use, but that is exactly the spirit.
We also understand that this is not the bond market or the large-cap stock market. It is a young market with many known imperfections. The data is neither as plentiful nor as high-quality as what you might find in those established markets.
But that is also where the opportunity lies: trying to make a difference by providing good data, as well as financial instruments and derivatives built on top of that data, which can become useful tools for industry stakeholders.
Ben Lorica. To give our audience a sense of how good the data is, I’m sure there are things you’re not comfortable sharing, but can you describe your data sources at a high level?
Steve Hou. Our data sources are a mixture of publicly available list prices and proprietary transaction prices obtained from data partners, cloud providers, and our sister exchange.
Silicon Data has a sister company called Compute Exchange. As the name suggests, Compute Exchange allows people to transact compute—whether they have capacity to sell or are looking to acquire it.
We take the totality of the data we observe and normalize it. Let me use an analogy. Suppose you wanted to create an index for unfurnished one-bedroom apartment rentals in New York City. How would you do it?
You would collect rental agreements from across the city, but they would come in different shapes and forms. Some apartments would be furnished and others unfurnished. Some would be rented for three, six, nine, or 12 months. They would also have idiosyncratic differences involving location and other characteristics.
You would take all those attributes, normalize them so the properties could be compared on an apples-to-apples basis, and consolidate the results into a single index. That index would represent the general price dynamics and trends for renting an unfurnished one-bedroom apartment in New York. It would not necessarily be the precise price you would be quoted when you went out to rent a particular apartment.
The same principle applies to GPUs.
I would actually love to share a chart of our indices while we discuss this. I have a small deck that I recently prepared for some Wall Street audiences. Here, we’re showing four major indices: the H100, the B200, which is Blackwell, the older A100, and the H200, which is a variant of the H100 with greater memory capacity and bandwidth.
We calculate the indices for these chips using the same methodology I just described. We spend a lot of time qualifying data sources to make sure they have been around long enough and are reliable. We also make sure there is sufficient market coverage before calculating an index for a particular chip.
People have asked us about the GB300, which is the latest and hottest Blackwell chip. We do not yet have enough source data, so we have not launched an index for it. As the data becomes more available and plentiful, we will launch one.
Ben Lorica. You can stop sharing and we can keep talking. How frequently is the index refreshed? As a quant, am I looking at intraday or hourly data?
Steve Hou. It is a daily calculated index.
Ben Lorica. There aren’t many traders in this space who require intraday data at this point, so daily is good enough for most people, right?
Steve Hou. This is a daily reference benchmark rate. The market is not yet exchange-traded, and there aren’t many meaningful intraday ticks. A daily update is a reasonable representation of the market.
Ben Lorica. All right. Here’s the first big-picture question, Steve. I don’t know whether you heard yesterday, but Cerebras and AMD have announced a new partnership.
Let’s talk about AMD first. NVIDIA’s market share is obviously overwhelming right now, but are there indicators showing that things are beginning to go well for AMD?
Steve Hou. It certainly seems so. AMD is also one of the chip providers we track. We try to track everything.
Ben Lorica. The reason I’m asking you is exactly that: A lot of people can offer opinions, but you have data.
Steve Hou. Right. At the moment, there has not been a lot of deployment. But we believe the inference market will eventually be much larger than the training market.
The way the industry is progressing—especially with efficient allocation and the disaggregation of inference into prefill and generation, along with memory management—aligns well with AMD’s vision and history. AMD has a long history of working across GPUs, CPUs, and related systems.
Lisa Su is tremendously talented. The partnership with Cerebras is similar to NVIDIA’s tie-up with Groq and, more recently, Etched, another startup effort. OpenAI has also been adopting related approaches.
Everyone is recognizing that the demand for inference will be enormous, and they want greater inference efficiency. There probably isn’t a market for a thousand different types of custom inference ASICs, but there will probably be room for a few. These partnerships make a lot of sense. I personally published something on what AMD is doing.
Ben Lorica. You said that you track AMD GPUs. Are there signals in the data showing a positive trend for AMD?
Steve Hou. Not just yet. What we’re seeing is encouraging news at the margin, but deployment takes time. Scaling up systems, constructing data centers, and installing the hardware involve lengthy lead times.
The transition from a market dominated by training toward one increasingly dominated—and eventually perhaps overwhelmingly dominated—by inference will also take time.
We don’t see the signals in the deployment data just yet, but there is every reason to expect that they will eventually appear.
Ben Lorica. What about Cerebras? Do you track Cerebras?
Steve Hou. I don’t know a great deal about the company. I like the founder, and I think he’s a nice guy.
It’s remarkable that Cerebras started as an effort to build a training chip, became very good at training, and then stumbled upon this massive inference use case. It seems promising.
Ben Lorica. Because Cerebras has its own private cloud and its systems aren’t openly bought and sold, you don’t track it.
Steve Hou. Not yet. I don’t think it has been very widely deployed outside its own environment.
Ben Lorica. It has a private cloud focused on inference. I’m friends with both the CEO and the person who runs the private cloud.
The issue for Cerebras right now—and I don’t know the exact numbers—is that a large portion of its revenue comes from OpenAI. It needs to diversify away from that.
Steve Hou. We should track it. At the moment, I don’t think we do, but it is absolutely something we should consider.
Ben Lorica. One more obscure hardware question: Do you track TPUs?
Steve Hou. We’re actively talking to Google at the moment.
The way we look at the market is that we try to track, as much as possible, how chips are being used by neo-clouds and other pure AI compute providers, whether they offer bare metal or another form of access. These providers represent the marginal, purest form of AI compute.
If you look at rental rates for AI compute from hyperscalers, the same chips tend to be priced at a significant premium because of bundling and other attributes.
Right now, TPUs are not widely available outside Google, at least as far as we know. There has been news that Google is beginning to supply or support data centers running on TPUs. But TPUs have been in such short supply, even for Google’s internal use, that we have not seen many used elsewhere.
Eventually, we absolutely want to track them. As a matter of principle, we want to track everything we can.
Ben Lorica. With the advent of Chinese open-weight models, it seems as though token prices are going down. There appears to be price pressure on Anthropic and OpenAI. Is that a fair assessment?
People are also starting to realize that token pricing alone isn’t necessarily the right metric. What matters is the cost of completing a task. Some models might be cheaper per token, but they can spin around in circles and consume many more tokens.
Nevertheless, it seems that open-weight models are creating a downward trend in prices. Are you seeing that in the data?
Steve Hou. There is an intuition that prices are under pressure, but we have not yet observed the published prices themselves going down.
For each new generation of models, supply still appears constrained and outmatched by demand. We have not seen the token prices for those models decline.
That said, the effective expenditure per million tokens among active API users is trending down. We publish a token index based on the sample we observe across public inference and routing platforms.
We combine input and output token prices with the volume associated with each model to produce an aggregate, expenditure-weighted price index. That index is coming down.
In other words, the behavior of price-sensitive users shows substitution between models. The total volume of tokens is still growing secularly, but a larger share of the marginal growth is shifting toward cheaper models.
A few months ago, everyone was trying to use Claude and other frontier models as much as possible, or perhaps Codex. Now the mix is changing.
But, as you mentioned, token price is not the same as task price. If a cheaper model costs one-tenth as much but I have to rerun it many times, never get the answer I want, give up, and switch to a frontier model, I haven’t saved money. I might have spent more than if I had successfully one-shotted the task with the frontier model.
We don’t yet have an index based on capability or actual task-level intelligence. That is a work in progress, but measuring task capability is difficult.
Ben Lorica. It would also be useful if you could track changes in usage patterns over time, although I think that will be difficult.
Suppose I’m J.P. Morgan. Twelve months ago, I used OpenAI and Claude 50% of the time each. Now, my usage of those models has fallen to 10% each because 80% of the time I’m using open-weight models. Most of my tasks are routine. I still use the high-end models for the most complex and challenging tasks, but my mix of model usage has changed.
That isn’t something you can determine from your data, right?
Steve Hou. Not directly. We don’t observe individual usage data. We observe activity at the model level and aggregate it to the market level.
However, you can assume that different users face similar incentives when optimizing their budgets. We all make trade-offs between price and quality when we shop.
Those incentives may be similar across J.P. Morgan, smaller companies, and individual users, even if the degree or speed of substitution differs. If the general direction is similar, what we measure is probably indicative of the broader trend.
Indeed, our data has often led the market narrative in terms of where trends are headed. That is how we think about the question.
Ben Lorica. I’m actually a user of OpenCode and OpenRouter. OpenRouter periodically publishes a “state of the market” report covering usage patterns.
But many people go to OpenRouter to use open-weight models. If I’m primarily using Claude or OpenAI, I might simply call those endpoints directly.
To some extent, reports about the state of the LLM market from places like OpenRouter don’t provide a complete picture of what people are using.
Steve Hou. First, OpenRouter is not our only source. But generally speaking, the available data does not let us observe traffic directly on the frontier labs’ own platforms. We also don’t observe all the traffic going through the hyperscalers. We are talking to some of them, but we don’t have direct visibility.
That is why I mentioned the similarity of incentives. Suppose I work at a major enterprise. Let’s use an example—
Ben Lorica. Bloomberg.
Steve Hou. Exactly. At Bloomberg, you’re allowed to use certain models. You cannot substitute freely because you can use only the models that have been made available to you.
You might be able to substitute among different model tiers within an approved provider. Claude, for example, offers different tiers such as Haiku and Opus. But the potential substitution is more limited.
As companies adopt more types of models, you will see more substitution. In general, when people are given a choice, they make rational trade-offs.
I don’t necessarily agree that someone who uses an open inference-routing platform is interested only in open models.
Ben Lorica. I use Claude and OpenAI through OpenRouter as well. But if I decide that I’m going to use mainly Claude for a particular task, I might call Claude directly because OpenRouter takes a cut.
Steve Hou. Yeah.
If you imagine the overall market as a large square, the data we observe is not a vertical slice of the square, because that would be a representative sample. Instead, we observe one corner of the square because of the limitations of the available data.
But that corner is still indicative. It tells us the direction of travel when people are making quality-versus-cost substitutions.
Suppose frontier models become increasingly powerful, the capability gap grows, and users get more value for their money because the task price is lower with frontier models. It should not be surprising if the data begins moving in that direction.
Between March and June, our token index increased very strongly. People were using more frontier models during that period. The trend is probably more meaningful than the index’s absolute level.
Ben Lorica. Right now, you’re mainly tracking inference. You aren’t tracking post-training directly.
There are services offering different forms of post-training, from supervised fine-tuning to reinforcement fine-tuning. That also uses compute, and it provides information about what people are doing.
If companies are performing post-training, they’re likely using open-weight models and building specialized models for narrow needs. That would suggest that they are moving away from the idea that one monolithic model will solve all their problems.
Are you paying attention to what people are doing in post-training?
Steve Hou. When you say post-training, I presume you mean companies such as Baseten—companies that help enterprises and other organizations with these activities.
Ben Lorica. There are companies like that, and there are also straightforward services. I can use fine-tuning as a service.
I haven’t seen much reinforcement fine-tuning as a service, but supervised fine-tuning as a service is certainly available.
Steve Hou. We don’t track it directly, but we track its net effect. We track GPU rental prices, so if a lot of demand is coming from post-training—
Ben Lorica. The problem is that after I perform post-training, I receive an endpoint. I start calling that endpoint, and you might not know that I’m using it because it is a specialized model I customized.
I did some post-training and now have an endpoint for my own model.
Steve Hou. But you’re still performing inference somewhere, using some type of infrastructure.
Ben Lorica. I might still be using Fireworks or another service.
Steve Hou. Fair enough, but somebody is ultimately using a GPU somewhere.
Ben Lorica. Exactly. But as far as model usage is concerned, you don’t know that I’m using this specialized model because you’re only tracking publicly visible models.
Steve Hou. Correct. Those models are not publicly available. If you’re asking whether our token index tracks those types of models, the answer is no. We don’t track private models.
Ben Lorica. Let’s return to GPU hardware and NVIDIA. Give us a quick, high-level snapshot of the state of the market for NVIDIA GPUs. What is happening with resale and rental values?
Steve Hou. The overall market is still tightening. When people look at rental prices, they often focus on the H100, the widely deployed Hopper workhorse.
I like to look at the A100, which is an aging chip from roughly five years ago.
Ben Lorica. Yeah, yeah.
Steve Hou. Its price is still going up. It is an aging chip that people might assume should have fallen out of favor, but it is fully booked. That is highly indicative of very strong inference demand.
Training workflows can be shifted around. As the B200 becomes more plentiful, some workflows move from the H100 to the B200. You can therefore see some movement at the front of the curve, in the on-demand market.
But then there is the forward curve, which we mentioned earlier. It tracks how GPU rental rates differ for contracts of different lengths. Let me quickly show you what it looks like.
If you’re familiar with finance, a commodity curve is often downward-sloping. Return to the rental analogy I used earlier: If you rent a New York apartment for 12 months, the daily rate will be cheaper than if you rent it for three days through Airbnb.
A longer-term tenant saves the landlord the cost of tenant turnover. The landlord doesn’t have to find another tenant, and the rental price could change during that period. In return for removing those risks, the long-term tenant receives a discount.
That is also what we typically see with GPUs. The blue line here represents November rental rates for H100 contracts of different lengths. The curve is clearly backwardated, which is the financial term for a downward-sloping term curve.
Over the four months through the end of March, the entire curve shifted upward. Every contract length became more expensive.
More importantly, the curve became less backwardated and even moved slightly toward contango at the longer end. In other words, it became less downward-sloping and slightly upward-sloping for longer-term contracts.
Cloud providers are comfortable not locking in customers or offering large discounts to long-term tenants. Instead, they are willing to roll over short-term contracts and take the opportunity to raise prices when those contracts renew.
Looking at the more recent dynamics, the purple line represents June 25. At the one-year term—the length at which much of the serious workload takes place—the rate was nearly unchanged between March and June.
Since June, however, we have seen a monotonic increase in the curve. The July 2 and July 14 curves overlap, and the green line for July 20 rises again.
Throughout this period, multiple providers have successively raised prices. That tells us the market is tight. It is Economics 101: Demand increases, supply doesn’t catch up, and the price adjusts.
We also use the term curve to infer the residual or fair value of outstanding chips. Here, we have the A100, H200, and B200. We calibrate our residual-value model against actual transaction values observed through our sister exchange, as well as other anecdotal numbers. Based on what we’re seeing, it is reasonably accurate.
The green bars for the A100, for example, show that it depreciated through most of 2025. Then, in late 2025, agentic AI activity picked up and the rental rate for the A100 began rising sharply.
The increase in rental rates was large enough to offset the normal time decay from depreciation. As a result, the chip’s estimated fair value has held roughly constant at approximately $5,000 throughout 2026.
The H100 actually appreciated during this period, as did the B200.
Despite all the discussion about Meta leasing compute and the possibility of excess capacity, we do not see evidence of excess in the fundamentals. Demand remains very strong, indicating a continuing compute shortage.
Of course, the stock market prices second derivatives. If the growth rate slows from 30% year over year to 25%, while the valuation is already pricing in 35%, the stock can pull back. But the underlying fundamentals still indicate a shortage.
Ben Lorica. Let me ask a bigger question. You can stop sharing for a moment.
One of the mysteries is that the AI data-center buildout is so heavily debt-financed. That is raising all sorts of alarm bells. Fundamentally, I’m not even sure that AI data centers are such a great business.
Look at neo-clouds such as CoreWeave and Nebius. It isn’t as though they’re producing profits left and right.
The demand may be high, but the profits aren’t there. You can use Claude or OpenAI for $20 a month and hammer those services. They probably start losing money on you within the first few days.
Something seems out of balance between the cost of providing these services and the amount of money companies can charge for them.
This also returns us to where we started. The models themselves are becoming increasingly commoditized, placing downward pressure on prices. There seems to be a mismatch.
Steve Hou. Let me unpack some of that. This is a very new industry.
At the beginning, the $20-per-month pricing was essentially a teaser price introduced by OpenAI when usage patterns were almost entirely based on chat queries. That was before agentic usage emerged.
Today, billing for frontier models is moving increasingly toward usage-based or API-based pricing. You are charged based on how much you use.
The alternative is an all-you-can-eat, fixed monthly price. Those plans are increasingly being throttled or repriced. Prices are also rising.
Ben Lorica. But can these companies actually make money on inference? What does inference cost in terms of energy and GPU compute, and how much are they charging?
It is hard to perform that calculation. All we can say is that no one is profitable. Obviously, these companies have other expenses, but it is difficult to raise prices much further because customers now have many more options.
Steve Hou. Let me separate a couple of issues. First, the margins on inference itself are very high. Various sources have confirmed that. Anthropic reportedly has gross margins of approximately 70% to 80% on some of its inference business.
Ben Lorica. But how long can it maintain that? The models are getting better.
Steve Hou. That is why I need to step back. This is ultimately about predictions. Everyone is expressing a view, and the market reflects a debate among those different opinions.
Nebius, for example, is a neo-cloud company whose stock price is much higher today than it was exactly a year ago. We do not know what the future holds or what the eventual business model will look like.
There is a consensus view that this is a bad-unit-economics business built around rapidly depreciating hardware assets. That remains to be seen. The operators of these data centers will have to prove whether this can become a better business.
Ben Lorica. The other issue is the use of debt. Everyone is using special-purpose vehicles, which can keep some of the debt off their balance sheets. When they report earnings, they can say they don’t have as much debt—
Steve Hou. I don’t have as much to say about the financing decisions. I want to separate the things we can observe from the things we cannot.
Imagine that all new investment paused today and no new models were trained. Serving the existing models would still be very—
Ben Lorica. I agree.
Steve Hou. The question becomes whether, with today’s model capabilities, we can still expect further growth in adoption. I think the answer is probably yes. Adoption is still very limited.
Taken as a whole, these companies are spending heavily because they are continuing to build. The model labs are investing in the training of new models, and data-center operators are constructing new data centers. As a result, their overall cash flow is going to be—
Ben Lorica. Fueled by debt.
Steve Hou. Fueled by debt. Capital expenditures are typically financed with debt.
Ben Lorica. But there is a significant difference here. The debt is now being held by regular people through pension funds and insurance companies. It isn’t just held by private equity anymore.
If there is a crash, it could cascade through the system. It would not simply be a matter of one company going bankrupt.
I’m bullish about AI. I’m worried about the financial structure, not the long-term prospects of the technology. But many things about the financial side concern me.
Steve Hou. Isn’t there a famous saying that the stock market is always climbing a wall of worry?
It is a genuine concern. It is part of the risk.
Ben Lorica. There is also the fact that the models are becoming increasingly commoditized.
Another thing the market underappreciates is that the infrastructure for post-training is becoming simpler. The entire post-training stack will become much more accessible to ordinary companies, allowing them to create specialized models themselves.
Many startups are working on different elements of post-training. As the tooling becomes simpler, more AI usage will shift toward specialized models that companies can build themselves.
That doesn’t necessarily hurt what you’re doing, Steve, because AI usage will still grow. Those specialized models will still use GPUs.
Steve Hou. I’m an economist, so I tend to think in terms of the eventual equilibrium. Some of these partial-equilibrium effects can move in particular directions.
Overall, I think we will still have frontier models and less-than-frontier models—
Ben Lorica. I’m not saying otherwise. I’m saying post-training will shift more of the workload away from general-purpose frontier models and toward specialized models. I can see the tools becoming simpler.
Steve Hou. That is a conjecture. We don’t know how much of the workload will shift, but some degree of shifting will occur.
Ben Lorica. It will shift because enterprises want to own that compounding loop.
Steve Hou. I don’t disagree. Enterprise AI adoption will involve multimodel workflows.
The question is whether the overall amount of AI usage will become much greater. Suppose I double my AI token usage—
Ben Lorica. AI usage will continue to grow, I think.
But this is an industry that relies on a few companies spending enormous amounts of money on data centers. There is a financial dimension separate from the technology.
The AI technology will grow, but there is also a financial bubble. There is no question about that. The question is how we achieve a safe landing once that bubble—
Steve Hou. I don’t have a great deal of insight into that. But as an economist, I think we are going through a phase change and an evolution.
Right now, a large share of AI token demand is being funneled through the top two labs. I read somewhere that approximately 50% of AI compute demand is being committed through OpenAI and Anthropic, and Anthropic apparently cannot find enough capacity.
Eventually, one would hope that the market moves toward much broader enterprise adoption of multimodel AI. Demand would be funneled through many more enterprises, which would finance capital expenditures through operating expenditures.
Ben Lorica. If that happens, which is likely, you still have the circular-financing issue of neo-clouds relying heavily on those two frontier labs.
I think there will be somewhat of a hard landing, but hopefully we can work our way through it. That doesn’t change the fact that the technology will continue to be used.
We need to be careful and aware of the risks. I lived through the dot-com bust and the financial crisis, so I’m more wary of bubbles than some younger people.
The thing about bubbles is that once they burst, the market crashes, but people go back to building the next day.
The reason I bring up post-training and the shift away from frontier models is that it may accelerate this unwinding.
Steve Hou. Quite possibly. We don’t know exactly how it will work out. But you seem to have a very concrete thesis.
Ben Lorica. I’m looking at signals that seem to indicate that we’re headed toward some kind of correction. But that doesn’t change the fact that AI will keep moving forward.
The development of AI is being accompanied by a financial bubble, and financial bubbles cannot continue indefinitely. That seems to be what happens.
In closing, Steve, are there trends that remain below the radar but are beginning to appear in your data? One thing I don’t know whether you track is the rise of Chinese GPUs.
Steve Hou. Huawei certainly has chips in development. You saw the most recent interview with the DeepSeek CEO. The suggestion is that these systems could eventually become important, although they aren’t there yet.
Ben Lorica. I recently read an article suggesting that the Chinese government’s attitude is essentially, “If you use NVIDIA, you’re a traitor.” They’re trying to jump-start their domestic chip industry.
But you aren’t tracking that yet. You aren’t seeing it in the data.
Steve Hou. For one thing, there is very limited visibility into what China is doing. We don’t know how much training and inference are still being performed with NVIDIA chips versus what Huawei may be able to offer in the future.
When you ask about trends that are below the radar, I don’t think people are paying enough attention to what the Chinese hardware ecosystem may become capable of.
There is also a presumption that Chinese enterprise demand for AI will not emerge because Chinese enterprises have historically been reluctant to pay for SaaS software. But they are increasingly finding value in cloud computing and AI.
It is possible that China will develop its own AI capital-expenditure cycle now that its models have moved closer to the frontier.
There are developments occurring below the radar in China, although it isn’t yet clear what direction they will take.
Something surprising could emerge on the hardware side. Huawei may produce much more capable and powerful GPUs. However, when it comes to token throughput and efficiency, I don’t know how they will compare with the latest NVIDIA hardware. NVIDIA has a very powerful hardware ecosystem.
Chinese demand could also experience its own version of the Jevons paradox, which we haven’t really seen yet.
Ben Lorica. How global is your data? We obviously have AI demand in China, the United States, and Europe, but it is increasingly emerging in the developing world, Southeast Asia, and the Gulf states.
Is your data global in that sense? Is GPU pricing like the price of oil, with a global spot price?
Steve Hou. We’re trying to create a global benchmark, although regional differences obviously exist. We also calculate regional prices.
China has its own distinct ecosystem. Outside China, however, there is more of a general global market.
Ben Lorica. Where does your data show the greatest demand by region? Obviously, the United States is first. What comes after the United States?
Steve Hou. Europe, followed by parts of the Asia-Pacific region. The fastest growth is actually occurring outside the United States, although the U.S. market is also growing very quickly.
Ben Lorica. India is now the most populous country. Is it beginning to appear in your data?
Steve Hou. Not to a significant extent. Activity could be growing at a fast rate, but the total quantity we observe is still very limited.
Ben Lorica. In closing, you also track RAM. I was speaking with someone from South Korea yesterday. Some of these memory companies had been left for dead several years ago. Now, you read stories about employees buying houses with their bonuses.
Will the RAM boom continue step by step with the GPU boom for the foreseeable future?
Steve Hou. For the foreseeable future, yes. Supply is not keeping up with demand.
I don’t know a great deal about the memory industry beyond that, but gross margins of approximately 80% don’t seem sustainable to me—or to anyone.
At some point, more supply will come online, demand will become more efficient in its use of memory, or both.
Ben Lorica. We’ve discussed AMD and the Cerebras ASIC. Is there any other up-and-coming hardware?
Intel always seems to have some type of GPU that never quite takes off. If you had to predict, 12 months from now, how much of your dataset will be driven by AMD?
Steve Hou. I can’t predict that.
Ben Lorica. Fair enough.
Steve Hou. That’s too difficult. But you saw the news about a new chip company called Etched. I think it has developed an incredible piece of technology.
The general trend is toward highly efficient inference chips designed to address the enormous size of the future inference market.
Ben Lorica. Everything also depends on TSMC. Even if companies want to manufacture more chips, there might not be sufficient manufacturing capacity.
Steve Hou. That’s right. At the moment, installed capacity is being constrained in many fairly mundane ways.
It is difficult to find enough power in the United States. Chips may not even be the primary constraint anymore. We don’t have a shortage of chips so much as a shortage of raw power. Where do you find it?
Ben Lorica. There is also growing local resistance to data centers for valid reasons involving energy use, noise, and other effects.
Steve Hou. We expect more fragmentation and deepening of the market. Smaller clusters will be able to serve post-trained, specialized models and combine them with frontier models.
The market won’t consist solely of gigawatt-scale data centers. There will be smaller clusters and 20- or 30-megawatt data centers powered by renewable energy and located closer to where the compute is being used.
For Wall Street, for example, those facilities might be close to New York and New Jersey.
Ben Lorica. With that, thank you, Steve. Thank you for humoring me and my doomsday predictions.
Steve Hou. Hopefully it won’t be doomsday. Market volatility is perfectly normal.

