After DeepSeek's 1100% price hike and industry-wide increases up to 80%, who offers the best value?
🔥 DeepSeek V4 Pro output $1.98/M (off-peak) ✅ Qwen3.5-Flash only $0.07/M inputPrices in USD per million tokens, as of August 26, 2026. DeepSeek now uses peak/off-peak pricing (peak: weekdays 9:00-12:00, 14:00-18:00 Beijing time).
| Model | Input | Output | Total (1M+1M) | vs Qwen3.5-Flash |
|---|---|---|---|---|
| Qwen3.5-Flash 🏆 | $0.07 | $0.26 | $0.33 | 1x |
| DeepSeek V4 Flash (off-peak) | $0.22 | $0.66 | $0.88 | 2.7x |
| DeepSeek V4 Flash (peak) | $0.44 | $1.32 | $1.76 | 5.3x |
| DeepSeek V4 Pro (off-peak) | $0.66 | $1.98 | $2.64 | 8x |
| DeepSeek V4 Pro (peak) | $1.32 | $3.96 | $5.28 | 16x |
| Gemini 2.5 Flash | $0.30 | $2.50 | $2.80 | 8.5x |
| Gemini 3.5 Flash | $1.50 | $9.00 | $10.50 | 32x |
| Claude Sonnet 5 (current) | $2.00 | $10.00 | $12.00 | 36x |
| Claude Sonnet 5 (after Sept) | $3.00 | $15.00 | $18.00 | 55x |
| GPT-5.5 | $5.00 | $30.00 | $35.00 | 106x |
| GPT-5.5 Pro | $30.00 | $180.00 | $210.00 | 636x |
Sources: Official pricing pages, DeepSeek API docs, Morgan Stanley research, GitHub state-of-llm-apis, August 2026.
| Model | Input | Output | Context | Notes |
|---|---|---|---|---|
| Qwen3.5-Flash | ¥0.2 | ¥2 | 1M | Multimodal, cache hit ¥0.02 |
| qwen-turbo | ¥0.3 | ¥0.6 | 131K | Thinking mode output ¥3 |
| Qwen3.5-Plus | ¥0.8 | ¥4.8 | 128K | Enhanced reasoning |
| DeepSeek V4 Flash (off-peak) | ¥1.5 | ¥4.5 | — | Weekends all off-peak |
| DeepSeek V4 Flash (peak) | ¥3 | ¥9 | — | Up 200-350% |
| DeepSeek V4 Pro (off-peak) | ¥4.5 | ¥13.5 | — | Cache hit ¥0.15 |
| DeepSeek V4 Pro (peak) | ¥9 | ¥27 | — | Cache hit up 1100% |
Launched August 24, 2026, Alibaba Cloud's Wan3.0 generates up to 30-second videos natively and is the first to accept documents (doc/xls/ppt/pdf/md) as input.
| Resolution | Price/sec | Promo (Aug 24-Sep 23) | 30-sec video |
|---|---|---|---|
| 480p | $0.05 | $0.035 | $1.05 |
| 720p | $0.10 | $0.07 | $2.10 |
| 1080p | $0.20 | $0.14 | $4.20 |
Compared to Google Veo 3.1 at $0.40/sec, Wan3.0 1080p costs only 50% as much. New users get 50 seconds of free video generation.
Nvidia AI servers up 15%+, HBM supply shortages, DRAM up 90-95% in Q1. GPU rental rates: H100 from $1.70 to $2.35/GPU/hour.
GitHub's official announcement: "Agentic workflows have dramatically increased compute demands, with some single requests exceeding the cost of an entire plan." GitHub Copilot introduced session limits and 7-day token caps in April 2026.
OpenRouter data: DeepSeek-V4-Flash processed 11.31 trillion tokens in a single week (Aug 3-7). OpenCode platform processed 8 trillion tokens in a single day.
Morgan Stanley's report is titled "Farewell to Price Wars, Hello to Intelligence Wars." Tencent Cloud raised prices twice this year; Zhipu AI three times.
Qwen3.5-Flash supports 1M token context, multimodal input (text/image/video), Function Calling, and structured output. It scores 100/100 for pricing value in the coding category (lmmarketcap) and is 98% cheaper than the category average. It handles the vast majority of use cases.
Sign up for Alibaba Cloud and enter the Model Studio console — tokens are automatically credited, valid for 180 days. No credit card required for China region; international version offers $200 credit with a 30-day refund guarantee.
DeepSeek V4 Pro remains competitive for complex reasoning, but at peak pricing it approaches mid-tier international model costs. Use Qwen3.5-Flash daily and switch to V4 Pro only when you need its unique capabilities — preferably off-peak or on weekends.
Alibaba Cloud International offers $200 free credits (30-day refund guarantee), with Qwen3.5-Flash at $0.07/$0.26 per M tokens across Singapore, Germany, and US nodes.
New users: $200 cloud credit + 70M AI tokens + 80+ products, 30-day refund guarantee
Claim $200 Free Credit → ECS from $4.50/moChina users: Get 70M free tokens on Bailian →