Released · official changelog 2026-09-10

DeepSeek V4.1 Flash
Released: The Guide

DeepSeek shipped V4.1 Flash on 2026-09-10: the smallest model in its new architecture family, native multimodal, 1M context and 384K max output, with the new API id deepseek-flash. The legacy deepseek-v4-flash and -vision-exp ids are temporarily routed to V4.1 Flash. Everything here is anchored to the official changelog, release note and pricing page.

Updated 2026-09-21

Claude, GPT and Chinese models on one key, billed per token; check /models for callable models.

#DeepSeek V4.1 Flash#Released 2026-09-10#Official id deepseek-flash#1M / 384K

Four verifiable official facts

2026-09-10

Official release

The changelog went live with DeepSeek-V4.1-Flash that day, with native multimodal visual understanding.

deepseek-flash

Official model id

Release note and docs: set model to deepseek-flash to call V4.1 Flash.

1M / 384K

Context / max output

Pricing page lists CONTEXT LENGTH 1M, MAX OUTPUT 384K.

552B / 8B / 16B

MoE active parameters

Release note: a 552B-parameter MoE, 8B active for input, 16B for output.

What V4.1 Flash is

V4.1 Flash is the smallest model in DeepSeek's new architecture family, built for faster inference, higher throughput and native vision. The official API id is deepseek-flash; the previous deepseek-v4-flash and deepseek-v4-flash-vision-exp are retired and temporarily routed to V4.1 Flash, billed at the Flash price. On QCode, the listed id is deepseek-v4.1-flash and is callable directly.

The 09-10 release and the 09-14 routing detour

Officially released 2026-09-10. The release note first said all deepseek-v4-pro requests would route to V4.1-Flash at Flash rates starting 2026-09-14 04:00 UTC; the changelog's 2026-09-10 entry then corrected course—continuing V4 Pro with unchanged billing. Where the two official statements disagree, defer to the changelog's 2026-09-10 entry: it states 'In response to user demand, we have decided to continue…' (a new decision), and the release note's routing claim was not upheld. The pricing page records no decision—the 2026-09-20 capture still lists deepseek-v4-pro as a row with no retirement note.

Timeline

2026-09-09

This page's 2026-09-09 draft recorded that no official V4.1 entry existed yet—only media reports. The official release came the very next day: the clearest case for not scheduling migrations off a media timetable.

2026-09-10

The changelog released DeepSeek-V4.1-Flash on 2026-09-10 with API id deepseek-flash; legacy flash ids route temporarily and bill at the Flash price.

2026-09-14

The release note's planned 2026-09-14 v4-pro routing date; the changelog's 2026-09-10 entry then said it continues deepseek-v4-pro with unchanged billing. The model is still listed on the pricing page as of 2026-09-20.

Confirmed vs conflict

✅ Officially confirmed

Released 2026-09-10 (smallest new-architecture model, native vision); API id deepseek-flash; 1M context / 384K max output; legacy flash ids temporarily routed and billed at the Flash price; peak/off-peak pricing effective at release. Each maps to the official changelog, release note and pricing page.

⚠️ Read both sides

The v4-pro routing: the release note (route to Flash from 2026-09-14) conflicts with the changelog's 2026-09-10 entry (continue v4-pro, billing unchanged). Defer to the changelog's 2026-09-10 entry — the changelog's 2026-09-10 entry states 'In response to user demand, we have decided to continue…'. Nothing on the pricing page comments on v4-pro—its only footnotes cover the retired legacy flash names and off-peak rates. No leaked third-party benchmarks or prices are used.

How to read the routing conflict

Trust only the release note

It says v4-pro routes away on 2026-09-14, but shutting down a v4-pro integration on that line misses the changelog correction—the changelog's 2026-09-10 entry states a new decision to continue by user demand with unchanged billing. The pricing page carries no such note: the 2026-09-20 capture has footnotes only on the legacy flash names and on off-peak rates.

Defer to the changelog entry that states the decision

When two official statements conflict, take the document that actually records the decision: the changelog's 2026-09-10 entry says 'In response to user demand, we have decided to continue…' (a new decision), so v4-pro continues, billing unchanged. The pricing page only shows today's state—deepseek-v4-pro is still a listed row—and states no decision. Readers can confirm by checking the v4-pro unit price on their own bill.

How to wire it up on QCode

For new integrations use the /models-listed id deepseek-v4.1-flash: same endpoint, same key, just set the model field. Legacy requests pinned to deepseek-v4-flash are served by V4.1 Flash on DeepSeek's side, but we have no per-request evidence for which version answers that id on QCode, so prefer deepseek-v4.1-flash to pin the version explicitly.

On QCode

deepseek-v4.1-flash is on sale and listed on /models with real 30-day usage, sharing one key and one quota with Claude, GPT and the rest—switching is just the model field. The new official id deepseek-flash is not currently listed on /models; whether it can be set on QCode is decided by /models.

FAQ

Is DeepSeek V4.1 Flash released?

Yes. DeepSeek released DeepSeek-V4.1-Flash on the changelog on 2026-09-10, native multimodal, API id deepseek-flash.

What's the official id, and what do I use on QCode?

The new official id is deepseek-flash. On QCode, use the /models-listed id deepseek-v4.1-flash—it's on sale with real usage; the official deepseek-flash id is not currently listed on /models, so whether it can be set follows /models.

Context and max output?

The official pricing page lists 1M (1,000,000) context and 384K max output.

Does the old deepseek-v4-flash still work?

On DeepSeek's side it is temporarily routed to V4.1 Flash and billed at the Flash price; requests are still accepted. On QCode we have no per-request proof of which version answers that id, so switch new integrations to deepseek-v4.1-flash.

Is deepseek-v4-pro being re-routed?

Official sources disagree: the release note said v4-pro routes to Flash from 2026-09-14, then the changelog's 2026-09-10 entry states it continues v4-pro with unchanged billing. Defer to the changelog's 2026-09-10 entry (it is the only official text that records the decision); the pricing page still lists deepseek-v4-pro as of 2026-09-20, so check your own v4-pro unit price to confirm.

How is it priced?

Peak/off-peak: peak is UTC Mon–Fri 01:00–04:00 and 06:00–10:00, everything else off-peak at half the peak rate, with a lower cache-hit price. New pricing took effect 2026-09-10 04:00 UTC. Use the official pricing page's table for the day.

Sources

DeepSeek official changelog, release note news260910, and models/pricing page, all captured 2026-09-18. Where two official statements conflict (release note vs changelog) both are shown with which one wins. QCode ids and usage come from /models and the 30-day call statistics.

One key across the DeepSeek family

deepseek-v4.1-flash and deepseek-v4-pro share one QCode key and quota with Claude and GPT—switching is just the model id. Priced at official rate × service fee; the key is live right after payment.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.