Kimi K3 Is Officially Out
Confirmed vs rumor
As of 2026-10-07, Moonshot calls Kimi K3 its most capable model to date: officially released on 2026-07-16, priced at $3 / $15 per 1M tokens on the official API, and there is still no K4 on the official model list. This page keeps tracking what is officially confirmed versus what remains community claims.
Current status (updated 2026-10-07)
As of 2026-10-07: Kimi K3 officially launched on 2026-07-16, and its weights were open-sourced on 07-27 (2.8T Stable LatentMoE — the largest open-weight model ever). Moonshot's official model list describes kimi-k3 as Kimi's most capable model to date, with 2.8 trillion parameters, native visual understanding and a 1M-token context window. Official API pricing per 1M tokens: $3 input, $15 output, $0.30 cached input, with a 1,048,576-token context. The official platform changelog lists Kimi K3 as available on the Open Platform API in its July 2026 entry; its latest entry is September 2026, and Moonshot's newest Hugging Face release is still Kimi-K3. QCode /models already lists kimi-k3, with real calls over the last 30 days.
Confirmed vs community rumor
Officially confirmed
Release date 2026-07-16; weights open-sourced 07-27 (2.8T — the largest open weights ever); 1M-token context; official API price $3/$15; coding is one of the headline use cases. As of 2026-10-07 the official pricing page also lists $0.30 per 1M cached input tokens, with cache writes billed by TTL ($3 for 5 minutes, $6 for 1 hour).
Still not fully confirmed
Clarified: the parameter count (2.8T), the open-sourcing of weights on 07-27 and official API pricing ($3/$15) have all been published. The next-generation K4 has started training (multiple media reports on 07-29) — no date, no specs. Official check as of 2026-10-07: K4 does not appear in Moonshot's model list (English or Chinese), in the platform changelog (latest entry September 2026) or on its Hugging Face organization page. QCode bills per token; see /models for each model's price.
Why track Kimi K3
For Chinese developers, additional strong long-context options matter for agent workflows. Diversifying providers helps when any single quota window tightens.
How Kimi K3 fits in the 2026 landscape
Vs GLM-5.2
Different training, closed vs open weights trade-offs. Track both.
Vs GPT long context
Chinese ecosystem strength and cost profile may differ.
Provider hedge
Another option when Claude or Codex windows tighten.
Access via QCode today
kimi-k3 is callable on QCode with the same key as the other models on /models; billing is per token, with each model's price on /models. As of 2026-10-07 no Kimi K4 turns up on Moonshot's official model list, changelog or Hugging Face page, and QCode does not offer it.
Release date 2026-07-16; weights open-sourced 07-27 (2.8T — the largest open weights ever); 1M-token context; official API price $3/$15; coding is one of the headline use cases. As of 2026-10-07 the official pricing page also lists $0.30 per 1M cached input tokens, with cache writes billed by TTL ($3 for 5 minutes, $6 for 1 hour).
Clarified: the parameter count (2.8T), the open-sourcing of weights on 07-27 and official API pricing ($3/$15) have all been published. The next-generation K4 has started training (multiple media reports on 07-29) — no date, no specs. Official check as of 2026-10-07: K4 does not appear in Moonshot's model list (English or Chinese), in the platform changelog (latest entry September 2026) or on its Hugging Face organization page. QCode bills per token; see /models for each model's price.
For Chinese developers, additional strong long-context options matter for agent workflows. Diversifying providers helps when any single quota window tightens.
As of 2026-10-07: Kimi K3 officially launched on 2026-07-16, and its weights were open-sourced on 07-27 (2.8T Stable LatentMoE — the largest open-weight model ever). Moonshot's official model list describes kimi-k3 as Kimi's most capable model to date, with 2.8 trillion parameters, native visual understanding and a 1M-token context window. Official API pricing per 1M tokens: $3 input, $15 output, $0.30 cached input, with a 1,048,576-token context. The official platform changelog lists Kimi K3 as available on the Open Platform API in its July 2026 entry; its latest entry is September 2026, and Moonshot's newest Hugging Face release is still Kimi-K3. QCode /models already lists kimi-k3, with real calls over the last 30 days.
Release date 2026-07-16; weights open-sourced 07-27 (2.8T — the largest open weights ever); 1M-token context; official API price $3/$15; coding is one of the headline use cases. As of 2026-10-07 the official pricing page also lists $0.30 per 1M cached input tokens, with cache writes billed by TTL ($3 for 5 minutes, $6 for 1 hour).
Clarified: the parameter count (2.8T), the open-sourcing of weights on 07-27 and official API pricing ($3/$15) have all been published. The next-generation K4 has started training (multiple media reports on 07-29) — no date, no specs. Official check as of 2026-10-07: K4 does not appear in Moonshot's model list (English or Chinese), in the platform changelog (latest entry September 2026) or on its Hugging Face organization page. QCode bills per token; see /models for each model's price.
For Chinese developers, additional strong long-context options matter for agent workflows. Diversifying providers helps when any single quota window tightens.
As of 2026-10-07: Kimi K3 officially launched on 2026-07-16, and its weights were open-sourced on 07-27 (2.8T Stable LatentMoE — the largest open-weight model ever). Moonshot's official model list describes kimi-k3 as Kimi's most capable model to date, with 2.8 trillion parameters, native visual understanding and a 1M-token context window. Official API pricing per 1M tokens: $3 input, $15 output, $0.30 cached input, with a 1,048,576-token context. The official platform changelog lists Kimi K3 as available on the Open Platform API in its July 2026 entry; its latest entry is September 2026, and Moonshot's newest Hugging Face release is still Kimi-K3. QCode /models already lists kimi-k3, with real calls over the last 30 days.
Release date 2026-07-16; weights open-sourced 07-27 (2.8T — the largest open weights ever); 1M-token context; official API price $3/$15; coding is one of the headline use cases. As of 2026-10-07 the official pricing page also lists $0.30 per 1M cached input tokens, with cache writes billed by TTL ($3 for 5 minutes, $6 for 1 hour).
Clarified: the parameter count (2.8T), the open-sourcing of weights on 07-27 and official API pricing ($3/$15) have all been published. The next-generation K4 has started training (multiple media reports on 07-29) — no date, no specs. Official check as of 2026-10-07: K4 does not appear in Moonshot's model list (English or Chinese), in the platform changelog (latest entry September 2026) or on its Hugging Face organization page. QCode bills per token; see /models for each model's price.
Why track Kimi K3
For Chinese developers, additional strong long-context options matter for agent workflows. Diversifying providers helps when any single quota window tightens.
As of 2026-10-07: Kimi K3 officially launched on 2026-07-16, and its weights were open-sourced on 07-27 (2.8T Stable LatentMoE — the largest open-weight model ever). Moonshot's official model list describes kimi-k3 as Kimi's most capable model to date, with 2.8 trillion parameters, native visual understanding and a 1M-token context window. Official API pricing per 1M tokens: $3 input, $15 output, $0.30 cached input, with a 1,048,576-token context. The official platform changelog lists Kimi K3 as available on the Open Platform API in its July 2026 entry; its latest entry is September 2026, and Moonshot's newest Hugging Face release is still Kimi-K3. QCode /models already lists kimi-k3, with real calls over the last 30 days.
Release date 2026-07-16; weights open-sourced 07-27 (2.8T — the largest open weights ever); 1M-token context; official API price $3/$15; coding is one of the headline use cases. As of 2026-10-07 the official pricing page also lists $0.30 per 1M cached input tokens, with cache writes billed by TTL ($3 for 5 minutes, $6 for 1 hour).
Clarified: the parameter count (2.8T), the open-sourcing of weights on 07-27 and official API pricing ($3/$15) have all been published. The next-generation K4 has started training (multiple media reports on 07-29) — no date, no specs. Official check as of 2026-10-07: K4 does not appear in Moonshot's model list (English or Chinese), in the platform changelog (latest entry September 2026) or on its Hugging Face organization page. QCode bills per token; see /models for each model's price.
Kimi K3 FAQ
Is K3 publicly usable now?
Yes. Since 2026-07-16, logged-in users can use K3·Max and K3 Cluster·Max. For API details, follow Moonshot's official platform docs.
How is this different from GLM-5.2 or GPT long context?
Different providers, different training mixes and context handling. Having options lets you route or fallback when one has quota pressure.
Should I wait for K3 or use current options?
No need to wait. K3 is out and its official API price is public ($3 / $15 per 1M tokens as of 2026-10-07), so test it on your own tasks now; before production, compare cost and stability on a small slice of traffic. Current Claude and GPT tiers remain the safe baseline.
Will QCode support K3 when released?
Already supported. kimi-k3 is listed on /models with real calls over the last 30 days; pricing follows the live list.
Stay ready with multi-provider access
One QCode key for current models; new ones added when live.
Related
Kimi K3 vs Claude Opus 5
Overlapping benchmark ranges, a roughly 40% price gap, and how Moonshot positions the model itself.
GLM-5.2 deep dive
The strongest open-weight coding model — capabilities, benchmarks and caveats.
AI Model Radar 2026
Our own composite model comparison, useful to cross-reference against the AA Index.