AI Inference Is Getting Cheaper, But That Could Make AI More Expensive.

fiverr
fiverr


  • AI inference prices are plunging as model providers battle for cost leadership.
  • DeepSeek’s price hikes show the race to the bottom may already be reversing.
  • Cheaper AI could drive more usage, shifting value from models to chips and infrastructure.

The AI industry’s next battle isn’t about who has the smartest model. It’s about who can run a task the cheapest. That shift became obvious this month, as U.S. labs slashed prices to hold off cheaper Chinese rivals, only for the cheapest of those rivals, DeepSeek, to turn around and raise its own rates.

Why Cost, Not Just Capability, Now Decides The Market

Chinese open-weight models have been quietly taking over. Reports this year show Chinese providers now handle more than half of all US-routed traffic on OpenRouter, the platform many developers use to compare and access models, up from roughly a tenth a year earlier. 

The pitch is simple: near-frontier performance at a fraction of the price. That pressure forced OpenAI’s hand. 

On July 30, it cut input pricing for GPT-5.6 Luna by 80%, from $1 to $0.20 per million tokens, with output pricing dropping from $6 to $1.20. Jefferies-cited data from Silicon Data shows average enterprise inference costs falling from roughly $2.04 per million tokens in late May to around $1.16-$1.18 in August, a yearly low.

okex

DeepSeek’s About-Face

Meanwhile, the company that started the price war just raised its own prices by a lot. Starting August 16-17, DeepSeek introduced peak and off-peak billing for its V4-Pro and V4-Flash models, with increases ranging from 50% to over 1,100% depending on the model and time of day. 

V4-Pro output jumps from $0.87 to $3.96 per million tokens at peak (9 am-12 pm and 2 pm- 6 pm Beijing time), or $1.98 off-peak. V4-Flash output rises from $0.28 to $1.32 at peak, and to $0.66 off-peak. 

DeepSeek says the change is meant to spread demand more evenly across its infrastructure. Even after the hike, its rates remain below many rivals, but the gap with OpenAI’s discounted Luna model has narrowed sharply, and in some peak-hour comparisons, disappeared entirely.

The Jevons Paradox Angle

Falling prices don’t necessarily mean falling bills. When each task gets dramatically cheaper, it becomes economical to run far more of them: multi-step agents, automated workflows, and processes that weren’t worth automating before. 

Analysts have already noted that agentic workloads chaining together dozens of calls, once impractical, are now cheaper than a single call was weeks earlier. That’s the classic Jevons paradox: efficiency gains can expand total consumption faster than they shrink the price per unit. In other words, enterprise AI spending could keep climbing even as the per-task cost drops.

Who Actually Loses

If inference keeps moving toward commodity pricing, model developers competing on raw output face the biggest squeeze, especially as OpenAI and Anthropic pursue IPOs at valuations near $1 trillion. Thinner token margins make those valuations harder to justify through model sales alone.

Infrastructure looks more insulated. More tasks, agents, and inference calls still require chips, cloud capacity, and data centers, regardless of who wins the pricing war. 

Related: Gemini 3.7 Flash Could Make Crypto the Ultimate Testing Ground for AI Agents

Disclaimer: The information presented in this article is for informational and educational purposes only. The article does not constitute financial advice or advice of any kind. Coin Edition is not responsible for any losses incurred as a result of the utilization of content, products, or services mentioned. Readers are advised to exercise caution before taking any action related to the company.





Source link

Coinmama

Be the first to comment

Leave a Reply

Your email address will not be published.


*