---
title: "55 Trillion Tokens a Week: Why Chinese Models Lead the World on Usage, Not on Benchmarks"
date: 2026-09-01
category: Foundation Models
site: NeuroAI
canonical: https://neuroai.site/a/na-55-trillion-tokens-a-week
language: en
---

# 55 Trillion Tokens a Week: Why Chinese Models Lead the World on Usage, Not on Benchmarks

> OpenRouter logged 113 trillion tokens of global model calls in one week. Chinese models accounted for 55.16 trillion — up 36% week on week, and ahead of the US for eighteen consecutive weeks. The top three were all Chinese.

## Key takeaways

- OpenRouter reported **113 trillion tokens** of total global model usage in a single week. Models from Chinese developers accounted for **55.16 trillion**, up **36%** week on week.

- Chinese models have now out-called US models for **eighteen consecutive weeks**.

- The top three by usage were all Chinese: **Zhipu GLM-5.3-Flash**, **DeepSeek-V4-Flash**, and **Xiaomi MiMo-V2.5**. GLM-5.3-Flash, unmasked after appearing anonymously as "Ox Alpha," took global first place at **15.7 trillion tokens in a week**.

- On **31 August 2026** China's Ministry of Industry and Information Technology issued guidance calling for increased procurement of "foundation models, agents and tokens" — framing tokens as a utility to be bought in volume.

## What a token is, and why it measures work

A token is the smallest unit an AI processes. In English roughly one word; in Chinese, roughly one to three characters.

Every time you have AI write a paragraph, analyse an image or revise a contract, you spend tokens.

So a weekly usage figure, in plain terms, is: how many times did the world actually use AI this week.

Fifty-five trillion is hard to picture. At roughly 700 characters per thousand tokens, it is equivalent to 3.8 trillion characters processed in a week.

## Why the winner was the cheap model, not the best one

The usage leader is not the strongest model. It is the best value model.

GLM-5.3-Flash is the first natively multimodal release in the GLM-5 series — it handles images and video. It is priced at one-tenth of GLM-5.3, and for a promotional period at one-twentieth, and it was released open-weight.

DeepSeek's V4-Flash-Vision-Exp was open-sourced the same evening and scored **59.3%** on a real-world software engineering evaluation, ahead of a leading closed model. Capability that is not far behind, at a price near the floor — developers vote with their feet.

Behind that sits a policy signal. When a ministry writes "tokens" into a procurement document, it is treating inference the way a government treats water or electricity: something to be bought in bulk, at scale, as infrastructure. One research estimate puts China's compute-rental market above RMB 260 billion this year.

## Who cheapness actually helps

The effect on ordinary people is more direct than it looks.

**People running a side business benefit first.** Copywriting, translation, code — tasks that once felt expensive on a frontier model now run free or for pennies on an open Chinese model. Several hundred calls a day stops being a cost decision.

**People using AI to get things done benefit.** The voice assistant on your phone, your office suite, your bank's customer service chatbot — all of them can be re-based onto cheaper domestic models. When service cost falls, either the price falls or the quality rises.

**Students and founders benefit most.** Open weights mean anyone can download and deploy. The cost of a top-tier model goes to zero.

## What to do now

One: go and use the free tiers. OpenRouter and most model providers give daily free calls without a top-up. Move your routine tasks — weekly reports, translation, research — onto them and see what actually breaks.

Two: watch the software you already pay for. If it runs on expensive models, expect either a price cut or a migration announcement. Switch subscriptions when it makes sense.

Three: do not pay for "AI capability" prematurely. Validate the need with a free open-weight model first. Most people never reach the point where they need to upgrade.

The price war has only started. What needs to hurry is not your wallet — it is how you use it.

## Honest limitations

OpenRouter measures usage routed through its own platform, which skews toward developers and open-weight models; it is not a complete census of global AI consumption. Usage share measures adoption, not capability, and Chinese models leading on calls says nothing about frontier performance. The eighteen-week streak is a statistic from one aggregator at one point in time.

*Sources: OpenRouter weekly usage rankings; MIIT guidance dated 31 August 2026; public model releases and pricing from Zhipu, DeepSeek and Xiaomi. Information only.*

---

Published by NeuroAI (https://neuroai.site/) — https://neuroai.site/a/na-55-trillion-tokens-a-week
Free to quote with attribution and a link to the original.
