The code doesn't lie, but the pricing model does. DeepSeek V4’s recent price increase — from ¥6 to ¥9 per million input tokens at peak — isn't a move to squeeze more revenue. It's a tactical signal that the Chinese AI giant is shifting from a 'volume-at-all-costs' strategy to a capacity-optimization playbook. And the timing couldn't be more revealing: ZhiPu GLM-5.3 dropped its benchmark results within hours, claiming superiority in 7 out of 9 agent-focused tests.
This isn't just a price war. It's a disclosure of two fundamentally different architectural philosophies. One is betting on infrastructure efficiency (DeepSeek), the other on raw benchmark performance (ZhiPu). And for anyone building on these APIs — especially in the crypto-native, coding-agent space — the decision tree just got a lot more complex.

Context: The Landscape of Chinese AI API Pricing
Before the hike, DeepSeek V4 was the undisputed price leader in China’s large language model API market. Its ¥6/¥18 peak input/output pricing was aggressively low, undercutting even Alibaba’s Qwen and ByteDance’s Doubao. The strategy worked: DeepSeek captured a significant share of developer mindshare, particularly among cost-sensitive teams building coding agents, automated trading bots, and on-chain analysis tools.
ZhiPu, on the other hand, priced its GLM-4 at ¥10/¥30, positioning itself as a premium alternative with better enterprise support. But the gap was wide enough to keep most developers on DeepSeek. Then came the V4 price increase, and ZhiPu immediately responded by launching GLM-5.3 at ¥8/¥28 — a price point that matches DeepSeek V4-Pro almost exactly (¥9/¥27).
Volatility is just interest for the impatient. The market now faces a rare moment of parity: two Chinese AI models, priced within ¥1 of each other, both claiming leadership in the coding-agent segment. The difference? DeepSeek’s cache pricing is ¥0.15 per million tokens — a staggering 60x cheaper than its peak input price. ZhiPu’s cache? ¥2. That's a 13x gap.
Core Insight: The Real Battle is Infrastructure, Not Benchmarks
Let’s dig into the numbers. DeepSeek’s off-peak pricing (half-price during low-traffic hours) and ultra-cheap cache hits reveal a hidden layer of infrastructure optimization. The †¥0.15 cache price is not just a discount; it's a mathematical signal about the marginal cost of inference. For a well-optimized KV-cache system, the cost of serving a cache-hit token is dominated by memory bandwidth, not compute. DeepSeek has clearly invested in a highly efficient cache layer — likely using prefix caching with aggressive reuse — and is now weaponizing that advantage.
Compare this to ZhiPu’s cache pricing at ¥2. That’s 1/4 of its peak input price, a ratio that aligns with industry norms (OpenAI’s cache is typically 50% off). But against DeepSeek’s 1/60 ratio, it’s a non-starter. For any developer building high-frequency, high-reuse applications — think code completion, template-based queries, or on-chain event analysis — the lifetime cost of using GLM-5.3 will be significantly higher.
This is where the contrarian angle emerges. The benchmark race is a distraction. ZhiPu’s 7/9 win in agent benchmarks looks impressive, but the margin is razor-thin: most wins are within 2-4 points. On Terminal Bench 2.1, DeepSeek trails by only 0.3 points. On NL2Repo, DeepSeek leads. The real story is that DeepSeek’s infrastructure is a generation ahead in terms of operational efficiency, while ZhiPu is betting on benchmark supremacy to justify a higher price. But the price is now equal, and the cache gap is massive.
Contrarian Angle: The Price Hike Was a Capacity Signal, Not a Revenue Grab
Here’s the counter-intuitive take: DeepSeek raised prices not because they could, but because they had to. Inference during peak hours is expensive, and if their GPU utilization is near 100%, the only way to manage demand is to increase the price of peak tokens. The off-peak half-price is a clever way to shift non-urgent workloads to low-traffic windows, optimizing overall GPU utilization.

But the cache pricing tells a deeper story. At ¥0.15, DeepSeek is essentially giving away cache hits at near-zero margin. This is a classic loss-leader strategy: lock developers into an ecosystem where they optimize their applications for DeepSeek’s cache architecture. Once a developer builds a custom cache layer, prompt templates, or agent workflows that rely on DeepSeek’s prefix caching, the switching cost becomes astronomical. ZhiPu’s cache pricing may be a reflection of its own infrastructure immaturity — or a deliberate choice to avoid the loss-leader trap. Either way, the gap is a moat for DeepSeek.
Liquidity is a river, not a pond. The same logic applies to developer attention. DeepSeek’s aggressive cache pricing is a river that will carry a steady flow of high-repeat users. ZhiPu’s benchmark wins are a pond — impressive at first glance, but limited in depth.

Takeaway: The Winner Will Be Determined by Ecosystem, Not Benchmarks
For the blockchain community, this has direct implications. Coding agents are the backbone of automated DeFi strategies, on-chain analysis, and smart contract auditing. The choice of API provider will affect not just cost, but also latency, reliability, and the ability to handle high-frequency, repetitive tasks.
If you’re building a trading bot that queries an LLM 10,000 times a day for market sentiment analysis, DeepSeek’s cache pricing could reduce your API bill by 90% compared to ZhiPu. If you’re building a one-shot agent that solves complex, novel problems, ZhiPu’s marginal benchmark advantage might tip the scale. But the cache gap is structural. It’s not a feature that can be closed overnight — it requires years of infrastructure investment.
Floor sweeps happen; rug pulls are a choice. The choice here is clear: DeepSeek is betting on infrastructure efficiency to win the long game, while ZhiPu is betting on benchmark performance to capture the short-term narrative. My money is on the infrastructure play. The code doesn’t lie — and the cache pricing is the most honest signal in the room.