Price cut just one day after launch — $0.15/M input tokens, 89% lower training cost, coding benchmarks match DeepSeek V4 Pro
✅ Input $0.15/M tokens ✅ Output $0.47/M tokens ✅ Cache hit $0.016/M 🔥 67% cheaper than DeepSeek V4 FlashPricing via Alibaba Cloud Model Studio (international) and Bailian platform (China), per million tokens.
| Billing Item | Price (per 1M tokens) |
|---|---|
| Input | $0.15 |
| Output | $0.47 |
| Cache hit | $0.016 |
| Billing Item | Launch Price (Aug 26) | Current Price (Aug 27+) | Change |
|---|---|---|---|
| Input | ¥1.0 | ¥0.8 | ↓20% |
| Output | ¥3.0 | ¥2.7 | ↓10% |
| Cache hit | — | ¥0.1 | — |
August 2026 pricing for major LLM APIs, USD per million tokens.
| Model | Input | Output | Active Params | Context |
|---|---|---|---|---|
| Qwen3.8-Flash NEW | $0.15 | $0.47 | 6B (MoE) | 1M |
| Qwen3.5-Flash (≤128K) | $0.11 | $0.67 | — | 1M |
| Qwen3.7-Flash | $0.034 | $0.132 | — | 256K |
| Qwen3.8-Max | $2.00 | $6.00 | — | 1M |
| DeepSeek V4 Flash (peak) | $0.42 | $1.26 | 13B | 128K |
| DeepSeek V4 Flash (off-peak 50%) | $0.21 | $0.63 | 13B | 128K |
| DeepSeek V4 Pro (peak) | $1.26 | $3.78 | — | 128K |
| GLM-5.3-Flash | ~$0.15+ | ~$0.47+ | 18B (MoE) | — |
Qwen3.8-Flash (codenamed Qwen3.8-Flash-Next during development) is an early preview of the Qwen4 architecture, with four core innovations:
Qwen Sparse Attention (QSA) works alongside GDN: GDN compresses historical context while QSA selects key information. In high-cache-hit 1M-token scenarios, this delivers 7.6x faster prefill and 4.9x faster decode.
Splits the traditional single residual pathway into 4 parallel branches, dynamically gating information flow for better cross-layer communication and training stability.
Uses a refined Muon + AdamW hybrid strategy with refitted scaling laws, eliminating batch warm-up and significantly improving convergence efficiency and training throughput.
Per Qwen's official technical report, Qwen3.8-Flash excels across multiple evaluation suites:
| Benchmark | Domain | Performance |
|---|---|---|
| SWE-bench Pro | Agentic coding | 62.5, leading peers |
| CoWorkBench | Long-horizon office tasks | Beats DeepSeek V4 Flash |
| Toolathlon Verified | Real-world tool use | Matches Claude Opus 4.6 |
| MathVision | Visual math reasoning | Strong multimodal |
| AndroidWorld | Mobile agent tasks | Embodied intelligence |
| ERQA | Embodied reasoning | Multimodal understanding |
Across 14 evaluations, the base model (6B activated) achieved the best results in 8. The fine-tuned version shows even stronger performance in coding, agents, and multimodal tasks.
Access Qwen3.8-Flash via Model Studio with OpenAI-compatible API. New users get free credit. The model serves on QwenCloud with 1M context by default.
Claim $200 Free Credit →Available on Alibaba Cloud's Bailian platform with new user free token quota. First to receive the latest Qwen releases.
Bailian Console →The open-weight version, Qwen3.8-Flash-Next, is available on Hugging Face and ModelScope for download, fine-tuning, and local deployment. The production version with built-in tools and 1M default context is served via QwenCloud API.
Qwen3.8-Flash is a MoE model (125B total / 6B active) optimized for cost efficiency and high throughput. Qwen3.8-27B is a dense vision-language model (all 27B active) for deep reasoning at $0.42/$3.08 per million tokens. Choose Flash for scale, 27B for depth.
When a request hits the context cache (e.g., repeated system prompts, long document prefixes), cached input tokens are billed at $0.016/M instead of $0.15/M — a 90% discount. This dramatically reduces costs for agents, RAG, and multi-turn conversations.
Qwen3.8-Flash supports text and image input natively (input_modalities: ["text", "image"]), with video analysis in the production version. It handles chart analysis, document OCR, and visual math.
5,000 RPM (requests per minute) and 5,000,000 TPM (tokens per minute), suitable for enterprise-grade high-concurrency deployments.
$200 free credit + ECS from $4.50/month + open-weight models available
Claim $200 Free Credit → View ECS Plans