DeepSeek V4 Price Hike, Fully Explained
After peak/off-peak pricing landed
As of 2026-10-08 DeepSeek's official price table is still peak/off-peak: peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays; all other hours, including weekends, are off-peak at half price. The English table is in USD (deepseek-v4-pro output $3.96 peak / $1.98 off-peak per million tokens), and the Flash tier's model name is now deepseek-flash. The scheme started on 2026-08-17, when V4-Pro peak output rose to ¥27 per million tokens (about +350%, according to Caixin and guancha.cn).
Updated 2026-10-08
Prices change: on QCode, Claude, GPT and Chinese models are billed per token — live rates on /models.
- Billed per token — live rates on /models
- Pay by card (Visa / Mastercard / AMEX), Apple Pay, Google Pay or crypto
- Self-service: top up and activate instantly
Four key numbers
V4-Pro peak output increase
Peak output went from ¥6 to ¥27 per million tokens. Cache-hit input rose the most — up to 1100%.
V4-Flash max increase
In the 2026-08-17 change the lightweight tier went up in step (per Caixin, 08-21). With the 2026-09-10 V4.1-Flash release DeepSeek says API prices were reduced; as of 2026-10-08 the English table lists deepseek-flash off-peak at $0.15 input (cache miss) and $0.6 output per million tokens. Off-peak remains half of peak.
Peak windows (Beijing time)
Official wording: peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, excluding Chinese public holidays (9:00-12:00 and 14:00-18:00 Beijing time). All other hours, including weekends and Chinese public holidays in full, are off-peak (checked 2026-10-08).
Effective date
Console warning on 08-06 → official announcement on 08-13 (same day as V4 Pro GA) → effective at midnight 08-17 Beijing time, which the official change log gives as 16:00 UTC on August 16, 2026.
What actually changed
As of 2026-10-08 the DeepSeek API is still billed by time of use, with off-peak rates at half the peak rates. The scheme dates from 2026-08-17, when DeepSeek moved its API from flat pricing to peak/off-peak pricing: every billable item goes up during peak hours and returns to half price off-peak. The official framing is 'a return to value-based token pricing' — the 2.5x promotional discount from May and the low-price strategy are over, freeing compute budget to 'develop stronger large models.' The hike was announced the same day as V4 Pro GA (0813) and took effect within two weeks.
Market reaction
The price-rise note landed alongside “V4 Pro GA with much stronger agent capability”, read across the industry as DeepSeek moving from “price slayer” to value pricing. The table was rewritten twice after that: on 2026-09-10 new rates came in with V4.1-Flash and the Flash tier was cut, and the changelog's 2026-09-10 entry withdrew the V4-Pro redirect, keeping V4 Pro with billing unchanged. The developer conversation is still about re-running the numbers: off-peak scheduling plus cache-hit optimisation. The official price table as of 2026-10-08: the Flash tier's model name is deepseek-flash (version DeepSeek-V4.1-Flash); the legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but the corresponding models have been retired, and their requests are served by DeepSeek-V4.1-Flash at the Flash price. The English table is priced in USD, the Chinese one in CNY.
Timeline
Developer console warning: 'we plan a broad API price increase soon, expected to be significant.'
Official announcement: peak/off-peak pricing details + V4 Pro GA on the same day.
2026-08-17 the new rates take effect. On 2026-08-21 V4-Flash-Vision-Exp shipped at the V4-Flash price; on 2026-09-10 both it and deepseek-v4-flash were retired. The legacy names are still accepted and served by DeepSeek-V4.1-Flash at the Flash price; the official model name is now deepseek-flash.
Confirmed vs watch out
Confirmed
Verifiable word for word in DeepSeek's official docs (checked 2026-10-08): off-peak rates are half of peak; peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, excluding Chinese public holidays; the new prices took effect at 16:00 UTC on August 16, 2026; V4-Pro peak output is 27.0 CNY per million tokens in the Chinese table and $3.96 in the English one. The +350%/+1100%/+400% increase figures come from Caixin and guancha.cn reporting.
Watch out
The '11x price increase' headline refers to a single item — cache-hit input (¥0.025→¥0.3). Not everything went up 11x. Do the math on your own call profile; don't let the single largest item set the narrative.
How to cut costs after the hike
Off-peak scheduling
Move batch jobs, evals and nightly builds outside the peak windows: converting the official peak windows, weekdays 04:00-06:00 UTC and 10:00 UTC to 01:00 UTC the next day, plus weekends and Chinese public holidays in full, are off-peak at half the peak price.
Caching and model tiering
Reuse long prefixes to harvest the cache-hit discount; downgrade chore work to Flash or another cheap model, and reserve Pro for the hard stuff.
How to run your numbers
Three steps: ① pull your call profile (input/output/cache-hit ratios); ② re-price it under the new peak and off-peak rates; ③ push every deferrable workload into off-peak windows. For most batch-heavy workloads the adjusted total is flat or even lower.
On QCode
On QCode one key can call deepseek-v4.1-flash and deepseek-v4-pro, and the legacy name deepseek-v4-flash still works. Billing is per token; per-model prices are listed on /models. The peak/off-peak, CNY and USD figures on this page come from DeepSeek's own price table, not from QCode.
FAQ
When did DeepSeek raise prices?
Effective at midnight (Beijing time) on 2026-08-17, which the official change log gives as 16:00 UTC on August 16, 2026. The console warning came on 08-06 and the official announcement on 08-13, the same day as V4 Pro GA.
How much did prices rise?
For 2026-08-17: V4-Pro peak output ¥27 per million tokens; according to Caixin and guancha.cn, that was about +350%, with cache-hit input up to +1100% and V4-Flash up to +400%; off-peak was half the peak. With the 2026-09-10 V4.1-Flash release DeepSeek says API prices were reduced; as of 2026-10-08 the Flash tier's official model name is deepseek-flash (DeepSeek-V4.1-Flash), at $0.15 input (cache miss) and $0.6 output per million tokens off-peak.
How are peak and off-peak windows defined?
Per the official table as of 2026-10-08: peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays (9:00-12:00 and 14:00-18:00 Beijing time). All other hours, including weekends and Chinese public holidays in full, are off-peak.
Is DeepSeek still worth it after the hike?
As of 2026-10-08 the English table lists deepseek-v4-pro off-peak at $0.66 input (cache miss) and $1.98 output per million tokens, half the peak price. If you can schedule and use caching, the real-world cost increase is manageable; run the peak and off-peak numbers on your own traffic mix first.
Do QCode prices go up too?
QCode bills per token, and per-model prices are listed on /models. The peak/off-peak figures on this page come from DeepSeek's own price table and are not QCode prices; for QCode, go by /models.
Is the '11x increase' real?
The single largest item is real (cache-hit input ¥0.025→¥0.3), but most billable items rose far less. Price your own call profile, not the biggest single number.
Sources
DeepSeek API docs: Models & Pricing (English USD table and Chinese CNY table) and the Change Log entries of 2026-09-10 and 2026-08-13, fetched 2026-10-08. Historical increase figures: DeepSeek's pricing note (2026-08-13, reported by Xinhua Finance), Guancha (2026-08-17), Caixin (2026-08-21), Kaiyuan Securities (2026-08-06), Xueqiu summaries.
One key for DeepSeek V4 Pro and V4.1 Flash
QCode bills per token; per-model prices and the current model list are on /models.
Related reading
Claude Code gateway errors and official fixes
The advisor_20260301 400, Extra inputs are not permitted, rejected beta headers and 1M context: exact error text, cause and fixed version.
DeepSeek V5 tracker
The 'stronger model' rumors behind the price hike.
Prompt caching price tracker 2026
Cache read and write prices, TTLs and minimum cache lengths for Claude, GPT and DeepSeek, with two recomputable examples (as of 2026-10-07).
Not affiliated with DeepSeek. Checked 2026-10-08; prices and peak hours follow DeepSeek's official pages. Model availability follows /models.