Key takeaways
- OpenRouter reported 113 trillion tokens of total global model usage in a single week. Models from Chinese developers accounted for 55.16 trillion, up 36% week on week.
- Chinese models have now out-called US models for eighteen consecutive weeks.
- The top three by usage were all Chinese: Zhipu GLM-5.3-Flash, DeepSeek-V4-Flash, and Xiaomi MiMo-V2.5. GLM-5.3-Flash, unmasked after appearing anonymously as "Ox Alpha," took global first place at 15.7 trillion tokens in a week.
- On 31 August 2026 China's Ministry of Industry and Information Technology issued guidance calling for increased procurement of "foundation models, agents and tokens" — framing tokens as a utility to be bought in volume.
What a token is, and why it measures work
A token is the smallest unit an AI processes. In English roughly one word; in Chinese, roughly one to three characters.
Every time you have AI write a paragraph, analyse an image or revise a contract, you spend tokens.
So a weekly usage figure, in plain terms, is: how many times did the world actually use AI this week.
Fifty-five trillion is hard to picture. At roughly 700 characters per thousand tokens, it is equivalent to 3.8 trillion characters processed in a week.
Why the winner was the cheap model, not the best one
The usage leader is not the strongest model. It is the best value model.
GLM-5.3-Flash is the first natively multimodal release in the GLM-5 series — it handles images and video. It is priced at one-tenth of GLM-5.3, and for a promotional period at one-twentieth, and it was released open-weight.
DeepSeek's V4-Flash-Vision-Exp was open-sourced the same evening and scored 59.3% on a real-world software engineering evaluation, ahead of a leading closed model. Capability that is not far behind, at a price near the floor — developers vote with their feet.
Behind that sits a policy signal. When a ministry writes "tokens" into a procurement document, it is treating inference the way a government treats water or electricity: something to be bought in bulk, at scale, as infrastructure. One research estimate puts China's compute-rental market above RMB 260 billion this year.
Who cheapness actually helps
The effect on ordinary people is more direct than it looks.
People running a side business benefit first. Copywriting, translation, code — tasks that once felt expensive on a frontier model now run free or for pennies on an open Chinese model. Several hundred calls a day stops being a cost decision.
People using AI to get things done benefit. The voice assistant on your phone, your office suite, your bank's customer service chatbot — all of them can be re-based onto cheaper domestic models. When service cost falls, either the price falls or the quality rises.
Students and founders benefit most. Open weights mean anyone can download and deploy. The cost of a top-tier model goes to zero.
What to do now
One: go and use the free tiers. OpenRouter and most model providers give daily free calls without a top-up. Move your routine tasks — weekly reports, translation, research — onto them and see what actually breaks.
Two: watch the software you already pay for. If it runs on expensive models, expect either a price cut or a migration announcement. Switch subscriptions when it makes sense.
Three: do not pay for "AI capability" prematurely. Validate the need with a free open-weight model first. Most people never reach the point where they need to upgrade.
The price war has only started. What needs to hurry is not your wallet — it is how you use it.
Honest limitations
OpenRouter measures usage routed through its own platform, which skews toward developers and open-weight models; it is not a complete census of global AI consumption. Usage share measures adoption, not capability, and Chinese models leading on calls says nothing about frontier performance. The eighteen-week streak is a statistic from one aggregator at one point in time.
Sources: OpenRouter weekly usage rankings; MIIT guidance dated 31 August 2026; public model releases and pricing from Zhipu, DeepSeek and Xiaomi. Information only.
