Flash-tier comparison · both on the list

GLM-5.3-Flash vs DeepSeek V4 Flash
List $0.15/$0.50 vs $0.14/$0.28; both 1M windows, MIT weights

Zhipu shipped GLM-5.3-Flash on 2026-08-26 (320B-A18B, 1M, MIT). Official list $0.15 / $0.50 per million; launch half-off $0.075 / $0.25 until 2026-09-09 24:00 UTC+8. DeepSeek's live Flash tier today is V4.1-Flash from 2026-09-10, off-peak list $0.15 / $0.6, window 1M / 384K; the 0731 build is retired and its rate at the time was $0.14 / $0.28. Both ids show usage > 0 here over 30 days — one key, one endpoint, switch the model field.

Updated 2026-09-20

Switch models on the same endpoint and measure the cost per token yourself; check /models for callable models.

#glm-5.3-flash#deepseek-v4-flash#$0.15/$0.50 vs $0.14/$0.28#both 1M

Four numbers you can quote

$0.15 / $0.50

GLM-5.3-Flash official list

In / out per million. Launch half-off is $0.075 / $0.25 through 2026-09-09 24:00 UTC+8. QCode follows official × fee; on 2026-08-27 the billed pair was $0.14 / $0.49.

$0.14 / $0.28

DeepSeek V4 Flash uncached / output

0731's official rates in 2026-08 (peak/off-peak from 2026-08-16 16:00 UTC). That build was retired on 2026-09-10; the live one is deepseek-v4.1-flash, window 1M in / 384K max out.

1M / 1M

Both pitch a million-token context

GLM-5.3-Flash official 1M. DeepSeek V4 Flash 1M / 384K. Long-repo jobs: watch the real cut, not only the marketing window.

MIT / MIT

Both ship open weights

GLM-5.3-Flash landed on HF on 2026-08-26. DeepSeek V4 Flash-0731 is MIT too, but was retired on 2026-09-10. Open weights ≠ you must self-host; the API path follows official list prices.

Flash versus Flash, not another GLM guide

The site already has a GLM-5.3-Flash guide, a DeepSeek V4 Flash 0731 history page, and GLM-5.2 vs DeepSeek V4 Pro. This page only lines up the cheap tiers on official price, window and license — and DeepSeek's live cheap tier is V4.1-Flash from 2026-09-10. GLM is a 320B-A18B multimodal MoE; 0731 was the same Flash scale with fresh post-training and a jumped agent score (historical). Pick by task, not slogan.

Why Flash-vs-Flash showed up after 08-26

Launch day ended the Ox Alpha stealth trial and swapped in z-ai/glm-5.3-flash. On the DeepSeek side: 0731 ran in preview from 2026-07-31, took peak/off-peak pricing on 2026-08-16, and was retired on 2026-09-10 in favour of V4.1-Flash. The two official output rates now sit around $0.50 against $0.6. After QCode added GLM-5.3-Flash on 2026-08-27, both ids take real traffic on one endpoint.

Timeline

2026-07-31

2026-07-31 API preview of DeepSeek V4 Flash 0731; the id stayed deepseek-v4-flash, and the official uncached / output rate then was $0.14 / $0.28. The build was retired on 2026-09-10.

2026-08-16 16:00 UTC

2026-08-16 DeepSeek Flash moves to peak/off-peak. Do not quote only the daytime list.

2026-08-26 / 2026-08-27

GLM-5.3-Flash ships and opens its weights; QCode adds it the next day (2026-08-27) at $0.14 / $0.49 (official × fee). Official half-off runs through 2026-09-09 24:00 UTC+8.

Confirmed vs misread

Confirmed

Both official prices, 1M windows, MIT, GLM ship date 2026-08-26, discount end 2026-09-09 24:00 UTC+8, DeepSeek peak/off-peak from 2026-08-16, both ids usage > 0 here — vendor pages and already-shipped on-site copy, rechecked 2026-08-30.

Misread

“Flash means a laptop runs it” — GLM is 320B total; that is not a phone. “Output prices match” — official $0.50 vs $0.28. “QCode equals official half-off” — QCode is official × fee; do not merge the 2026-08-27 $0.14/$0.49 with $0.075/$0.25.

When to pick which

Lean GLM-5.3-Flash

Need vision in, the Zhipu tool chain, or the official half-off before 09-09. Architecture is 320B-A18B multimodal. Launch cache-read list $0.015. Trust the console for the bill.

Lean DeepSeek V4 Flash

Want the lower official output rate, peak/off-peak, or you already sit on the agent scores of DeepSeek's Flash line. Terminal Bench 2.1 public 82.7 belongs to the 0731 build and lives on its history page. Check 384K max out against the job.

How to switch on one endpoint (both sold)

The ids are glm-5.3-flash and deepseek-v4-flash. One already-provisioned key, one Chat Completions endpoint, change model. Estimate on official lists, then read the fee-multiplied number on your account. Do not treat the half-off window as a forever price.

Both Flash ids are callable on QCode

glm-5.3-flash and deepseek-v4.1-flash both show usage > 0 in our 30-day table, and you switch between them on the same endpoint with a single prepaid balance at official price × service rate. GLM listed $0.14/M in and $0.49/M out when it was added on 2026-08-27, with cache read cut to $0.04/M on 2026-09-08. The current official price of DeepSeek V4.1-Flash is $0.15/M input off-peak and $0.3/M peak, $0.6/M output off-peak and $1.2/M peak. The $0.14/$0.28 pair under the legacy name deepseek-v4-flash is the 2026-07-31 generation.

FAQ

Which table is the official price?

GLM: Zhipu / Z.ai $0.15/$0.50, half-off $0.075/$0.25 until 2026-09-09 24:00 UTC+8. DeepSeek: the 0731-era uncached / output $0.14/$0.28 plus peak/off-peak. QCode bill = official × fee.

Are both windows 1M?

GLM official 1M. DeepSeek Flash 1M in, 384K max out. Long-output jobs: check 384K first.

How is this different from GLM-5.2 vs DeepSeek V4 Pro?

That page is flagship / Pro. This page is Flash only. Do not copy Pro scores onto the Flash row.

Self-host pick?

Both MIT. GLM 320B-A18B VRAM is not something the word “Flash” summarizes. The 0731 build had its own homelab threads too, but it was retired on 2026-09-10. This page is API selection, not a rack list.

Can I try both on QCode?

Yes. Both ids have real usage. Same key, change model. Prices follow upstream.

Is GLM still half-off after 09-09?

The official discount is written through 2026-09-09 24:00 UTC+8. After that, list $0.15/$0.50 unless Zhipu posts otherwise. Do not freeze the discount as permanent.

Sources

GLM list, half-off window, 320B-A18B, 1M, MIT: Zhipu / Z.ai release of 2026-08-26 plus the on-site glm-5-3-flash guide and pricing page (rechecked 2026-08-30). DeepSeek's live $0.15 / $0.6 and 1M/384K: the official pricing page (captured 2026-09-18). The historical 0731 rates of $0.14 / $0.28 and peak/off-peak from 2026-08-16: the on-site deepseek-v4-flash-0731 history page. Retirement and temporary routing: the 2026-09-10 release note and pricing footnote (1). Usage: the CRS 30-day call-volume table retrieved 2026-09-18.

Both Flash models are sold; change model on the same endpoint

Official price × service fee. GLM half-off has an end date; DeepSeek has peak/off-peak. Estimate, then shift traffic.

Related

Prices and windows follow Zhipu / DeepSeek official pages and your bill. Discounts and peak/off-peak move. QCode bills official × fee and does not promise to match official half-off to the cent.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.