56.7T vs 16.5T tokens per week, four of the global top five, and Tencent's Hy4 taking #1. Token volume is money voting — here's how enterprises capture the "high-intelligence, low-cost" dividend.
🔥 56.7T tokens/week ✅ 19 weeks at #1 ✅ Model Studio from $0.15/M input 🚀 1M free tokens per modelOpenRouter is one of the world's largest LLM API aggregators, and its weekly usage is widely watched as the bellwether of real AI consumption. Last week's key figures:
| Metric | This week | WoW |
|---|---|---|
| Global total usage | 115T tokens | +1.77% |
| Chinese models | 56.7T tokens | +2.83% |
| US models | 16.5T tokens | -3.1% |
| China : US ratio | ~3.4 : 1 | 19th straight week ahead |
| Chinese models in global top 5 | 4 of 5 | Hy4 / GLM / DeepSeek ×2 |
| # | Model | Vendor · Country | Weekly tokens | WoW |
|---|---|---|---|---|
| 1 | Hunyuan Hy4 preview | Tencent · China | 14.7T | +379% |
| 2 | GPT-5.6 Luna | OpenAI · US | 12.9T | +66% |
| 3 | GLM-5.3 Flash | Zhipu · China | 12.4T | +101% |
| 4 | DeepSeek-V4-Flash (official) | DeepSeek · China | 12.4T | — |
| 5 | DeepSeek-V4-Flash (preview) | DeepSeek · China | 5.19T | — |
| 6 | MiniMax M3 | MiniMax · China | 5.02T | +95% |
Xiaomi MiMo-V2.5 and Gemini 3.7 Flash dropped out of the list this week. Source: OpenRouter, via public reporting (Sep 7, 2026).
The week's biggest story is Tencent Hunyuan's Hy4 preview: released and open-sourced on August 28, it took the #1 spot on OpenRouter within a week at 14.7T tokens (+379% WoW). Key specs:
| Dimension | Specification |
|---|---|
| Architecture | MoE, 770B total / 49B activated parameters |
| Context | 1M tokens (~1.5M Chinese characters) |
| License | Apache 2.0 open weights (HuggingFace / GitHub / ModelScope) |
| Focus | Agents, coding, productivity; generates playable game prototypes from one prompt |
| Blind eval | 163 experts, 203 engineering tasks, avg 2.99/4.00 — ahead of GLM-5.3 and Kimi K3 |
| Lightweight build | Sep 1: weights compressed from 1.5TB to ~214GB via in-house Sherry ternary quantization, lowering self-hosting barriers |
GLM-5.3 Flash, DeepSeek-V4-Flash and MiniMax M3 likewise crowded the leaderboard — all following the same playbook: open weights, rock-bottom API prices, and agent/coding-specific optimization. China's rise is a cluster breakthrough, not a single hit.
Behind the usage surge is a fact that favors every business: top-tier model APIs are now cheap enough to use liberally. The question is no longer "can we afford AI" but "which channel is the most stable, cost-effective, and compliant".
For teams in China and globally, Alibaba Cloud Model Studio (Bailian) is the simplest unified entry point — one platform, one API key, multiple leading models:
| Capability | Details |
|---|---|
| Model aggregation | Qwen3.8 family, DeepSeek-V4, MiniMax and more — text, image, video, voice; switch on demand |
| API compatibility | OpenAI- and Anthropic-compatible endpoints — migrate existing code with near-zero changes |
| Tool ecosystem | Works directly with Qwen Code, Claude Code, Qoder, OpenClaw and other leading AI coding/agent tools |
| New-user credits | Free tokens per model (90-day validity) |
| Cost controls | "Stop when free quota runs out" switch; alerts at 20% remaining and at exhaustion — no surprise bills |
| Enterprise-ready | Multi-tenant isolation, no queueing at peak; no training on your conversation data; contracts, invoices, SLA |
| Model / item | Input | Output | Cached input |
|---|---|---|---|
| Qwen3.8-Flash | $0.15 /M tokens | $0.47 | $0.016 |
| Context / limits | 262K native (1M via YaRN); 5,000 RPM / 5M TPM | ||
| Free credits | Per-model free quota for new accounts; check console for eligibility | ||
Figures from Alibaba Cloud Model Studio official pages (Sep 2026). The model price war is intense and adjustments are frequent — always confirm live pricing on the official campaign page.
Token volume reflects real-world usage scale and value, not universal technical supremacy. Chinese models clearly lead on open weights, price, agent/coding optimization, and Chinese-language scenarios — developers vote with their budgets. Frontier closed models still excel at some hard-reasoning tasks. The pragmatic enterprise approach is multi-model routing by task.
No. Model Studio aggregates Qwen, DeepSeek, MiniMax and other leading models across text, image, video, and voice — one API key, unified metering, switch on demand, with OpenAI/Anthropic-compatible endpoints.
Individuals and small teams should start with pay-as-you-go plus free credits — costs are minimal. Team plans suit organizations needing seat management, usage analytics, and budget caps, with admins assigning/recycling seats and monitoring per-member usage.
For most teams, the API wins — no GPUs to buy, no ops, pay-per-use, new models instantly available. Self-hosting open weights (Hy4, DeepSeek, Qwen) on Alibaba Cloud GPU instances makes sense only with hard data-residency requirements or volumes large enough to amortize GPU costs.
Free credits for new users · OpenAI/Anthropic-compatible · One key for Qwen, DeepSeek, MiniMax
Claim Free Credits → View ECS Plans