⚠️ Important Update: Effective August 17, 2026, DeepSeek implemented peak/off-peak pricing. V4-Pro and V4-Flash rates changed significantly. Peak hours (weekdays 9:00-12:00, 14:00-18:00 Beijing time) cost 2x off-peak rates. This article reflects the latest pricing.

Alibaba Cloud Bailian · DeepSeek Access

Enterprise-grade SLA with free credits for new users

View DeepSeek Plans → AIGC Zone →
👉 Looking for cloud servers? Alibaba Cloud Up to 90% Off | Tencent Cloud Hot Deals | All Deals

1. Latest DeepSeek Pricing (Effective Aug 17, 2026)

DeepSeek now offers three main model versions with time-based pricing:

ModelBilling ItemOff-PeakPeakUnit
V4-Pro
(Flagship)
Input (cache miss)$0.63$1.25per 1M tokens
Input (cache hit)$0.02$0.04per 1M tokens
Output$1.88$3.75per 1M tokens
V4-Flash
(Fast/Light)
Input (cache miss)$0.21$0.42per 1M tokens
Input (cache hit)$0.007$0.014per 1M tokens
Output$0.63$1.25per 1M tokens
V3
(Stable)
Input/OutputInput $0.14 / Output $0.56 (flat rate)per 1M tokens

Peak hours: Weekdays 9:00-12:00 and 14:00-18:00 (Beijing time, UTC+8). All other 17 hours on weekdays, plus weekends and public holidays, are off-peak (50% of peak price). Prices shown in USD at approximate 7.2 CNY/USD rate.

2. Platform Comparison

Below is a comparison of major platforms offering DeepSeek-V3 API access:

PlatformModelInputOutputSLARating
Alibaba Cloud BailianDeepSeek-V3$0.14/1M$0.56/1M99.9%★★★★★
DeepSeek-R1$0.56/1M$2.22/1M
Volcano EngineDeepSeek-V3$0.14/1M$0.56/1M99.9%★★★★☆
SiliconFlowDeepSeek-V3$0.13/1M$0.53/1M99.5%★★★★☆
DeepSeek OfficialDeepSeek-V3$0.14/1M$0.56/1MNone★★★☆☆

Alibaba Cloud Bailian DeepSeek-R1 pricing: Input $0.56/1M tokens, output $2.22/1M tokens (Beijing region). R1 excels at mathematical reasoning, coding, and logical inference tasks.

3. Cost-Saving Strategies with Peak/Off-Peak Pricing

4. Enterprise Selection Criteria

Start Using DeepSeek API

New Alibaba Cloud Bailian users receive free token credits

Claim Free Credits → AI Agent Platform →

FAQ

How are peak/off-peak hours defined?

Peak: weekdays 9:00-12:00 and 14:00-18:00 Beijing time (UTC+8). Off-peak: all other hours on weekdays, plus all weekends and Chinese public holidays. Off-peak rates are 50% of peak rates.

What is the difference between V3 and V4?

V3 is the stable version with flat-rate pricing ($0.14 input / $0.56 output per 1M tokens). V4-Pro is the flagship reasoning model with peak/off-peak pricing. V4-Flash is the lightweight, cost-effective option. Choose based on your task complexity and budget.

What is prompt caching?

When multiple requests share the same prefix (such as a system prompt), DeepSeek caches that content. Subsequent requests pay the cache-hit rate, which is dramatically cheaper. V4-Pro cache-hit input is only $0.02/1M off-peak.

How can I further reduce API costs?

1) Run non-urgent batch jobs during off-peak hours; 2) Purchase pre-paid resource packages; 3) Optimize prompts to maximize cache hits; 4) Use Flash instead of Pro for simple tasks; 5) Purchase through authorized channels for additional discounts.