NeuroAI NEUROAINEUROAI.SITE
ESC

Why Nearly Half of American AI Tokens Now Run on Chinese Open Models

On OpenRouter, the share of tokens US companies spend on Chinese AI models has sat above 30% every week since February 2026 and peaked at 46%. Engineers are not switching for ideology — they are switching because the models are good and roughly fifty times cheaper.

2026-09-03 · 728 words · NeuroAI
Why Nearly Half of American AI Tokens Now Run on Chinese Open Models

The most reliable measure of whether a technology is actually being used is not a benchmark. It is a bill.

In 2026, a growing share of the world's AI bills are being paid to run models built in China — including by American companies.

Key takeaways

  • Usage, not downloads: the share of tokens US companies consumed on Chinese AI models via OpenRouter has been above 30% every week since 8 February 2026, reaching as high as 46%. The trailing 12-month average was 11%; in the first half of 2025 it was 4.5%.
  • Open source is winning the default slot: on OpenRouter, the open-source share of tokens processed jumped from 34% in January 2026 to 65% in June 2026, with more than 500 organisations switching from proprietary to open models.
  • The price gap is enormous: DeepSeek V4 Flash has been listed at $0.09 per million input tokens and $0.18 per million output tokens, against roughly $5 and $30 for the leading US frontier model — about 55x on input and 166x on output.
  • The performance gap has nearly closed: Stanford's 2026 AI Index put the gap between top Chinese and American models at 2.7% on Code Arena as of March 2026, with the two countries repeatedly trading the lead.
  • Named adopters: Coinbase, DoorDash's co-founder, and the startup Lindy have all publicly described moving production work to Chinese models.

The switch, in the customers' own words

The pattern in public statements is remarkably consistent: not "Chinese models are better," but "Chinese models are good enough, and the arithmetic is undeniable."

  • Coinbase CEO Brian Armstrong wrote in June that the exchange was experimenting with defaulting to open-weight models such as GLM 5.2 and Kimi 2.7 through its internal LLM gateway, while still letting engineers pick the right model for each task — with AI spending reported to have fallen by roughly half.
  • Lindy, a San Francisco startup, moved 100% of its production traffic from Anthropic's Claude to DeepSeek V4. Its CEO said the switch saved millions of dollars and improved performance on many core use cases.
  • Andy Fang, co-founder of DoorDash, said the company now routes "lower-level work" to Kimi K2.6 and reserves US frontier models for the hardest tasks — a combination he said vastly outperformed the previous all-US setup at lower cost, as reported by the Financial Times.

What these have in common is that none of them are research experiments. They are production systems with uptime requirements.

Why the gap closed

Three things happened at once.

Capability. Chinese labs closed the practical gap in exactly the workloads enterprises care about — coding and agentic behaviour. Zhipu's GLM-5.2 was reported to have taken first place overall in the Code Arena blind leaderboard with over a million participating users. MiniMax M3, Qwen3-Max-Thinking and DeepSeek V4 each carved out differentiated strengths in long context, code generation and large-parameter inference.

Cost. Efficiency became a strategy rather than a constraint. A Chinese AI executive quoted by Chinese media framed it directly: great models can be built "not by stacking computing power, but by relying on efficiency and innovation" — an argument shaped by necessity under export controls and vindicated in the market.

Licensing. Chinese models have largely shipped under permissive licences that allow modification and commercial use. For companies outside the United States, that removes a specific fear: dependence on a supplier that can be switched off by an export rule.

The generational shift

The numbers are stark when placed in sequence. Chinese open-source models accounted for 41% of global large-model downloads over the past year on one major international open-source platform, and a joint MIT–platform report found that Chinese-developed open models overtook the United States in global downloads for the first time, taking first place worldwide.

Alibaba's Qwen passed 1 billion cumulative Hugging Face downloads in January 2026, and by March accounted for over half of global open-source model downloads — a reversal from roughly four years earlier, when US models held about 60% of that share.

The shift from 4.5% to 46% in about eighteen months is not a marketing story. It is a migration.

Usage figures are as reported by OpenRouter and cited by CNBC in July 2026; pricing as listed on OpenRouter in June 2026; adoption examples from company statements reported in the Financial Times and Chinese media.

More in “Foundation Models” → · Back to home · Markdown version