DeepSeek V4.1 Flash
Released: The Guide
DeepSeek shipped V4.1 Flash on 2026-09-10: the smallest model in its new architecture family, native multimodal, 1M context and 384K max output, with the new API id deepseek-flash. The legacy deepseek-v4-flash and -vision-exp ids are temporarily routed to V4.1 Flash. Everything here is anchored to the official changelog, release note and pricing page.
Updated 2026-09-21
Claude, GPT and Chinese models on one key, billed per token; check /models for callable models.
- Billed per token — live rates on /models
- Pay by card (Visa / Mastercard / AMEX), Apple Pay, Google Pay or crypto
- Self-service: top up and activate instantly
Four verifiable official facts
Official release
The changelog went live with DeepSeek-V4.1-Flash that day, with native multimodal visual understanding.
Official model id
Release note and docs: set model to deepseek-flash to call V4.1 Flash.
Context / max output
Pricing page lists CONTEXT LENGTH 1M, MAX OUTPUT 384K.
MoE active parameters
Release note: a 552B-parameter MoE, 8B active for input, 16B for output.
What V4.1 Flash is
V4.1 Flash is the smallest model in DeepSeek's new architecture family, built for faster inference, higher throughput and native vision. The official API id is deepseek-flash; the previous deepseek-v4-flash and deepseek-v4-flash-vision-exp are retired and temporarily routed to V4.1 Flash, billed at the Flash price. On QCode, the listed id is deepseek-v4.1-flash and is callable directly.
The 09-10 release and the 09-14 routing detour
Officially released 2026-09-10. The release note first said all deepseek-v4-pro requests would route to V4.1-Flash at Flash rates starting 2026-09-14 04:00 UTC; the changelog's 2026-09-10 entry then corrected course—continuing V4 Pro with unchanged billing. Where the two official statements disagree, defer to the changelog's 2026-09-10 entry: it states 'In response to user demand, we have decided to continue…' (a new decision), and the release note's routing claim was not upheld. The pricing page records no decision—the 2026-09-20 capture still lists deepseek-v4-pro as a row with no retirement note.
Timeline
This page's 2026-09-09 draft recorded that no official V4.1 entry existed yet—only media reports. The official release came the very next day: the clearest case for not scheduling migrations off a media timetable.
The changelog released DeepSeek-V4.1-Flash on 2026-09-10 with API id deepseek-flash; legacy flash ids route temporarily and bill at the Flash price.
The release note's planned 2026-09-14 v4-pro routing date; the changelog's 2026-09-10 entry then said it continues deepseek-v4-pro with unchanged billing. The model is still listed on the pricing page as of 2026-09-20.
Confirmed vs conflict
✅ Officially confirmed
Released 2026-09-10 (smallest new-architecture model, native vision); API id deepseek-flash; 1M context / 384K max output; legacy flash ids temporarily routed and billed at the Flash price; peak/off-peak pricing effective at release. Each maps to the official changelog, release note and pricing page.
⚠️ Read both sides
The v4-pro routing: the release note (route to Flash from 2026-09-14) conflicts with the changelog's 2026-09-10 entry (continue v4-pro, billing unchanged). Defer to the changelog's 2026-09-10 entry — the changelog's 2026-09-10 entry states 'In response to user demand, we have decided to continue…'. Nothing on the pricing page comments on v4-pro—its only footnotes cover the retired legacy flash names and off-peak rates. No leaked third-party benchmarks or prices are used.
How to read the routing conflict
Trust only the release note
It says v4-pro routes away on 2026-09-14, but shutting down a v4-pro integration on that line misses the changelog correction—the changelog's 2026-09-10 entry states a new decision to continue by user demand with unchanged billing. The pricing page carries no such note: the 2026-09-20 capture has footnotes only on the legacy flash names and on off-peak rates.
Defer to the changelog entry that states the decision
When two official statements conflict, take the document that actually records the decision: the changelog's 2026-09-10 entry says 'In response to user demand, we have decided to continue…' (a new decision), so v4-pro continues, billing unchanged. The pricing page only shows today's state—deepseek-v4-pro is still a listed row—and states no decision. Readers can confirm by checking the v4-pro unit price on their own bill.
How to wire it up on QCode
For new integrations use the /models-listed id deepseek-v4.1-flash: same endpoint, same key, just set the model field. Legacy requests pinned to deepseek-v4-flash are served by V4.1 Flash on DeepSeek's side, but we have no per-request evidence for which version answers that id on QCode, so prefer deepseek-v4.1-flash to pin the version explicitly.
On QCode
deepseek-v4.1-flash is on sale and listed on /models with real 30-day usage, sharing one key and one quota with Claude, GPT and the rest—switching is just the model field. The new official id deepseek-flash is not currently listed on /models; whether it can be set on QCode is decided by /models.
FAQ
Is DeepSeek V4.1 Flash released?
Yes. DeepSeek released DeepSeek-V4.1-Flash on the changelog on 2026-09-10, native multimodal, API id deepseek-flash.
What's the official id, and what do I use on QCode?
The new official id is deepseek-flash. On QCode, use the /models-listed id deepseek-v4.1-flash—it's on sale with real usage; the official deepseek-flash id is not currently listed on /models, so whether it can be set follows /models.
Context and max output?
The official pricing page lists 1M (1,000,000) context and 384K max output.
Does the old deepseek-v4-flash still work?
On DeepSeek's side it is temporarily routed to V4.1 Flash and billed at the Flash price; requests are still accepted. On QCode we have no per-request proof of which version answers that id, so switch new integrations to deepseek-v4.1-flash.
Is deepseek-v4-pro being re-routed?
Official sources disagree: the release note said v4-pro routes to Flash from 2026-09-14, then the changelog's 2026-09-10 entry states it continues v4-pro with unchanged billing. Defer to the changelog's 2026-09-10 entry (it is the only official text that records the decision); the pricing page still lists deepseek-v4-pro as of 2026-09-20, so check your own v4-pro unit price to confirm.
How is it priced?
Peak/off-peak: peak is UTC Mon–Fri 01:00–04:00 and 06:00–10:00, everything else off-peak at half the peak rate, with a lower cache-hit price. New pricing took effect 2026-09-10 04:00 UTC. Use the official pricing page's table for the day.
Sources
DeepSeek official changelog, release note news260910, and models/pricing page, all captured 2026-09-18. Where two official statements conflict (release note vs changelog) both are shown with which one wins. QCode ids and usage come from /models and the 30-day call statistics.
One key across the DeepSeek family
deepseek-v4.1-flash and deepseek-v4-pro share one QCode key and quota with Claude and GPT—switching is just the model id. Priced at official rate × service fee; the key is live right after payment.
Related
DeepSeek model id changes
One troubleshooting page for the new id, temporary routing and vision-exp retirement since 09-10.
DeepSeek V4 Pro 0813 guide
Positioning, pricing and setup for the current v4-pro.
DeepSeek V4 Flash 0731 history
Historical record of the retired 0731 build; new integrations: this page and the id-change note.
This page separates 'officially confirmed' from 'conflicting statements—read both sides'. DeepSeek may update at any time; defer to official docs. No leaked third-party benchmarks or prices are used.