Status tracker · shipped

Kimi K3 Is Officially Out
Confirmed vs rumor

As of 2026-10-07, Moonshot calls Kimi K3 its most capable model to date: officially released on 2026-07-16, priced at $3 / $15 per 1M tokens on the official API, and there is still no K4 on the official model list. This page keeps tracking what is officially confirmed versus what remains community claims.

Current status (updated 2026-10-07)

As of 2026-10-07: Kimi K3 officially launched on 2026-07-16, and its weights were open-sourced on 07-27 (2.8T Stable LatentMoE — the largest open-weight model ever). Moonshot's official model list describes kimi-k3 as Kimi's most capable model to date, with 2.8 trillion parameters, native visual understanding and a 1M-token context window. Official API pricing per 1M tokens: $3 input, $15 output, $0.30 cached input, with a 1,048,576-token context. The official platform changelog lists Kimi K3 as available on the Open Platform API in its July 2026 entry; its latest entry is September 2026, and Moonshot's newest Hugging Face release is still Kimi-K3. QCode /models already lists kimi-k3, with real calls over the last 30 days.

Confirmed vs community rumor

Officially confirmed

Release date 2026-07-16; weights open-sourced 07-27 (2.8T — the largest open weights ever); 1M-token context; official API price $3/$15; coding is one of the headline use cases. As of 2026-10-07 the official pricing page also lists $0.30 per 1M cached input tokens, with cache writes billed by TTL ($3 for 5 minutes, $6 for 1 hour).

Still not fully confirmed

Clarified: the parameter count (2.8T), the open-sourcing of weights on 07-27 and official API pricing ($3/$15) have all been published. The next-generation K4 has started training (multiple media reports on 07-29) — no date, no specs. Official check as of 2026-10-07: K4 does not appear in Moonshot's model list (English or Chinese), in the platform changelog (latest entry September 2026) or on its Hugging Face organization page. QCode bills per token; see /models for each model's price.

Why track Kimi K3

For Chinese developers, additional strong long-context options matter for agent workflows. Diversifying providers helps when any single quota window tightens.

How Kimi K3 fits in the 2026 landscape

Vs GLM-5.2

Different training, closed vs open weights trade-offs. Track both.

Vs GPT long context

Chinese ecosystem strength and cost profile may differ.

Provider hedge

Another option when Claude or Codex windows tighten.

Access via QCode today

kimi-k3 is callable on QCode with the same key as the other models on /models; billing is per token, with each model's price on /models. As of 2026-10-07 no Kimi K4 turns up on Moonshot's official model list, changelog or Hugging Face page, and QCode does not offer it.

Release date 2026-07-16; weights open-sourced 07-27 (2.8T — the largest open weights ever); 1M-token context; official API price $3/$15; coding is one of the headline use cases. As of 2026-10-07 the official pricing page also lists $0.30 per 1M cached input tokens, with cache writes billed by TTL ($3 for 5 minutes, $6 for 1 hour).

Clarified: the parameter count (2.8T), the open-sourcing of weights on 07-27 and official API pricing ($3/$15) have all been published. The next-generation K4 has started training (multiple media reports on 07-29) — no date, no specs. Official check as of 2026-10-07: K4 does not appear in Moonshot's model list (English or Chinese), in the platform changelog (latest entry September 2026) or on its Hugging Face organization page. QCode bills per token; see /models for each model's price.

For Chinese developers, additional strong long-context options matter for agent workflows. Diversifying providers helps when any single quota window tightens.

As of 2026-10-07: Kimi K3 officially launched on 2026-07-16, and its weights were open-sourced on 07-27 (2.8T Stable LatentMoE — the largest open-weight model ever). Moonshot's official model list describes kimi-k3 as Kimi's most capable model to date, with 2.8 trillion parameters, native visual understanding and a 1M-token context window. Official API pricing per 1M tokens: $3 input, $15 output, $0.30 cached input, with a 1,048,576-token context. The official platform changelog lists Kimi K3 as available on the Open Platform API in its July 2026 entry; its latest entry is September 2026, and Moonshot's newest Hugging Face release is still Kimi-K3. QCode /models already lists kimi-k3, with real calls over the last 30 days.

Release date 2026-07-16; weights open-sourced 07-27 (2.8T — the largest open weights ever); 1M-token context; official API price $3/$15; coding is one of the headline use cases. As of 2026-10-07 the official pricing page also lists $0.30 per 1M cached input tokens, with cache writes billed by TTL ($3 for 5 minutes, $6 for 1 hour).

Clarified: the parameter count (2.8T), the open-sourcing of weights on 07-27 and official API pricing ($3/$15) have all been published. The next-generation K4 has started training (multiple media reports on 07-29) — no date, no specs. Official check as of 2026-10-07: K4 does not appear in Moonshot's model list (English or Chinese), in the platform changelog (latest entry September 2026) or on its Hugging Face organization page. QCode bills per token; see /models for each model's price.

For Chinese developers, additional strong long-context options matter for agent workflows. Diversifying providers helps when any single quota window tightens.

As of 2026-10-07: Kimi K3 officially launched on 2026-07-16, and its weights were open-sourced on 07-27 (2.8T Stable LatentMoE — the largest open-weight model ever). Moonshot's official model list describes kimi-k3 as Kimi's most capable model to date, with 2.8 trillion parameters, native visual understanding and a 1M-token context window. Official API pricing per 1M tokens: $3 input, $15 output, $0.30 cached input, with a 1,048,576-token context. The official platform changelog lists Kimi K3 as available on the Open Platform API in its July 2026 entry; its latest entry is September 2026, and Moonshot's newest Hugging Face release is still Kimi-K3. QCode /models already lists kimi-k3, with real calls over the last 30 days.

Release date 2026-07-16; weights open-sourced 07-27 (2.8T — the largest open weights ever); 1M-token context; official API price $3/$15; coding is one of the headline use cases. As of 2026-10-07 the official pricing page also lists $0.30 per 1M cached input tokens, with cache writes billed by TTL ($3 for 5 minutes, $6 for 1 hour).

Clarified: the parameter count (2.8T), the open-sourcing of weights on 07-27 and official API pricing ($3/$15) have all been published. The next-generation K4 has started training (multiple media reports on 07-29) — no date, no specs. Official check as of 2026-10-07: K4 does not appear in Moonshot's model list (English or Chinese), in the platform changelog (latest entry September 2026) or on its Hugging Face organization page. QCode bills per token; see /models for each model's price.

Why track Kimi K3

For Chinese developers, additional strong long-context options matter for agent workflows. Diversifying providers helps when any single quota window tightens.

As of 2026-10-07: Kimi K3 officially launched on 2026-07-16, and its weights were open-sourced on 07-27 (2.8T Stable LatentMoE — the largest open-weight model ever). Moonshot's official model list describes kimi-k3 as Kimi's most capable model to date, with 2.8 trillion parameters, native visual understanding and a 1M-token context window. Official API pricing per 1M tokens: $3 input, $15 output, $0.30 cached input, with a 1,048,576-token context. The official platform changelog lists Kimi K3 as available on the Open Platform API in its July 2026 entry; its latest entry is September 2026, and Moonshot's newest Hugging Face release is still Kimi-K3. QCode /models already lists kimi-k3, with real calls over the last 30 days.

Release date 2026-07-16; weights open-sourced 07-27 (2.8T — the largest open weights ever); 1M-token context; official API price $3/$15; coding is one of the headline use cases. As of 2026-10-07 the official pricing page also lists $0.30 per 1M cached input tokens, with cache writes billed by TTL ($3 for 5 minutes, $6 for 1 hour).

Clarified: the parameter count (2.8T), the open-sourcing of weights on 07-27 and official API pricing ($3/$15) have all been published. The next-generation K4 has started training (multiple media reports on 07-29) — no date, no specs. Official check as of 2026-10-07: K4 does not appear in Moonshot's model list (English or Chinese), in the platform changelog (latest entry September 2026) or on its Hugging Face organization page. QCode bills per token; see /models for each model's price.

Kimi K3 FAQ

Is K3 publicly usable now?

Yes. Since 2026-07-16, logged-in users can use K3·Max and K3 Cluster·Max. For API details, follow Moonshot's official platform docs.

How is this different from GLM-5.2 or GPT long context?

Different providers, different training mixes and context handling. Having options lets you route or fallback when one has quota pressure.

Should I wait for K3 or use current options?

No need to wait. K3 is out and its official API price is public ($3 / $15 per 1M tokens as of 2026-10-07), so test it on your own tasks now; before production, compare cost and stability on a small slice of traffic. Current Claude and GPT tiers remain the safe baseline.

Will QCode support K3 when released?

Already supported. kimi-k3 is listed on /models with real calls over the last 30 days; pricing follows the live list.

Stay ready with multi-provider access

One QCode key for current models; new ones added when live.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.