GLM-5.3-Flash vs DeepSeek V4 Flash
List $0.15/$0.50 vs $0.14/$0.28; both 1M windows, MIT weights
Zhipu shipped GLM-5.3-Flash on 2026-08-26 (320B-A18B, 1M, MIT). Official list $0.15 / $0.50 per million; launch half-off $0.075 / $0.25 until 2026-09-09 24:00 UTC+8. DeepSeek's live Flash tier today is V4.1-Flash from 2026-09-10, off-peak list $0.15 / $0.6, window 1M / 384K; the 0731 build is retired and its rate at the time was $0.14 / $0.28. Both ids show usage > 0 here over 30 days — one key, one endpoint, switch the model field.
Updated 2026-09-20
Switch models on the same endpoint and measure the cost per token yourself; check /models for callable models.
- Billed per token — live rates on /models
- Pay by card (Visa / Mastercard / AMEX), Apple Pay, Google Pay or crypto
- Self-service: top up and activate instantly
Four numbers you can quote
GLM-5.3-Flash official list
In / out per million. Launch half-off is $0.075 / $0.25 through 2026-09-09 24:00 UTC+8. QCode follows official × fee; on 2026-08-27 the billed pair was $0.14 / $0.49.
DeepSeek V4 Flash uncached / output
0731's official rates in 2026-08 (peak/off-peak from 2026-08-16 16:00 UTC). That build was retired on 2026-09-10; the live one is deepseek-v4.1-flash, window 1M in / 384K max out.
Both pitch a million-token context
GLM-5.3-Flash official 1M. DeepSeek V4 Flash 1M / 384K. Long-repo jobs: watch the real cut, not only the marketing window.
Both ship open weights
GLM-5.3-Flash landed on HF on 2026-08-26. DeepSeek V4 Flash-0731 is MIT too, but was retired on 2026-09-10. Open weights ≠ you must self-host; the API path follows official list prices.
Flash versus Flash, not another GLM guide
The site already has a GLM-5.3-Flash guide, a DeepSeek V4 Flash 0731 history page, and GLM-5.2 vs DeepSeek V4 Pro. This page only lines up the cheap tiers on official price, window and license — and DeepSeek's live cheap tier is V4.1-Flash from 2026-09-10. GLM is a 320B-A18B multimodal MoE; 0731 was the same Flash scale with fresh post-training and a jumped agent score (historical). Pick by task, not slogan.
Why Flash-vs-Flash showed up after 08-26
Launch day ended the Ox Alpha stealth trial and swapped in z-ai/glm-5.3-flash. On the DeepSeek side: 0731 ran in preview from 2026-07-31, took peak/off-peak pricing on 2026-08-16, and was retired on 2026-09-10 in favour of V4.1-Flash. The two official output rates now sit around $0.50 against $0.6. After QCode added GLM-5.3-Flash on 2026-08-27, both ids take real traffic on one endpoint.
Timeline
2026-07-31 API preview of DeepSeek V4 Flash 0731; the id stayed deepseek-v4-flash, and the official uncached / output rate then was $0.14 / $0.28. The build was retired on 2026-09-10.
2026-08-16 DeepSeek Flash moves to peak/off-peak. Do not quote only the daytime list.
GLM-5.3-Flash ships and opens its weights; QCode adds it the next day (2026-08-27) at $0.14 / $0.49 (official × fee). Official half-off runs through 2026-09-09 24:00 UTC+8.
Confirmed vs misread
Confirmed
Both official prices, 1M windows, MIT, GLM ship date 2026-08-26, discount end 2026-09-09 24:00 UTC+8, DeepSeek peak/off-peak from 2026-08-16, both ids usage > 0 here — vendor pages and already-shipped on-site copy, rechecked 2026-08-30.
Misread
“Flash means a laptop runs it” — GLM is 320B total; that is not a phone. “Output prices match” — official $0.50 vs $0.28. “QCode equals official half-off” — QCode is official × fee; do not merge the 2026-08-27 $0.14/$0.49 with $0.075/$0.25.
When to pick which
Lean GLM-5.3-Flash
Need vision in, the Zhipu tool chain, or the official half-off before 09-09. Architecture is 320B-A18B multimodal. Launch cache-read list $0.015. Trust the console for the bill.
Lean DeepSeek V4 Flash
Want the lower official output rate, peak/off-peak, or you already sit on the agent scores of DeepSeek's Flash line. Terminal Bench 2.1 public 82.7 belongs to the 0731 build and lives on its history page. Check 384K max out against the job.
How to switch on one endpoint (both sold)
The ids are glm-5.3-flash and deepseek-v4-flash. One already-provisioned key, one Chat Completions endpoint, change model. Estimate on official lists, then read the fee-multiplied number on your account. Do not treat the half-off window as a forever price.
Both Flash ids are callable on QCode
glm-5.3-flash and deepseek-v4.1-flash both show usage > 0 in our 30-day table, and you switch between them on the same endpoint with a single prepaid balance at official price × service rate. GLM listed $0.14/M in and $0.49/M out when it was added on 2026-08-27, with cache read cut to $0.04/M on 2026-09-08. The current official price of DeepSeek V4.1-Flash is $0.15/M input off-peak and $0.3/M peak, $0.6/M output off-peak and $1.2/M peak. The $0.14/$0.28 pair under the legacy name deepseek-v4-flash is the 2026-07-31 generation.
FAQ
Which table is the official price?
GLM: Zhipu / Z.ai $0.15/$0.50, half-off $0.075/$0.25 until 2026-09-09 24:00 UTC+8. DeepSeek: the 0731-era uncached / output $0.14/$0.28 plus peak/off-peak. QCode bill = official × fee.
Are both windows 1M?
GLM official 1M. DeepSeek Flash 1M in, 384K max out. Long-output jobs: check 384K first.
How is this different from GLM-5.2 vs DeepSeek V4 Pro?
That page is flagship / Pro. This page is Flash only. Do not copy Pro scores onto the Flash row.
Self-host pick?
Both MIT. GLM 320B-A18B VRAM is not something the word “Flash” summarizes. The 0731 build had its own homelab threads too, but it was retired on 2026-09-10. This page is API selection, not a rack list.
Can I try both on QCode?
Yes. Both ids have real usage. Same key, change model. Prices follow upstream.
Is GLM still half-off after 09-09?
The official discount is written through 2026-09-09 24:00 UTC+8. After that, list $0.15/$0.50 unless Zhipu posts otherwise. Do not freeze the discount as permanent.
Sources
GLM list, half-off window, 320B-A18B, 1M, MIT: Zhipu / Z.ai release of 2026-08-26 plus the on-site glm-5-3-flash guide and pricing page (rechecked 2026-08-30). DeepSeek's live $0.15 / $0.6 and 1M/384K: the official pricing page (captured 2026-09-18). The historical 0731 rates of $0.14 / $0.28 and peak/off-peak from 2026-08-16: the on-site deepseek-v4-flash-0731 history page. Retirement and temporary routing: the 2026-09-10 release note and pricing footnote (1). Usage: the CRS 30-day call-volume table retrieved 2026-09-18.
Both Flash models are sold; change model on the same endpoint
Official price × service fee. GLM half-off has an end date; DeepSeek has peak/off-peak. Estimate, then shift traffic.
Related
GLM-5.3-Flash guide
320B-A18B, Ox Alpha trial, vision — not expanded here.
DeepSeek V4 Flash 0731 (historical)
Post-training, peak/off-peak, TB2.1 score.
GLM-5.2 vs DeepSeek V4 Pro
Flagship comparison; do not mix with the Flash row.
Prices and windows follow Zhipu / DeepSeek official pages and your bill. Discounts and peak/off-peak move. QCode bills official × fee and does not promise to match official half-off to the cent.