Codex source & official docs · 2026-10-07

Codex error "Selected model is at capacity. Please try a different model": it is not your quota, and it is not a 429
Capacity, quota and 429 are three different things

As of 2026-10-07, this is Codex's capacity message — not exhausted quota and not a 429. In the source of OpenAI's official Codex repository (rust-v0.160.1, released 2026-10-05), “Selected model is at capacity. Please try a different model.” maps to the server error code server_is_overloaded; exhausted quota is a different line, “Quota exceeded. Check your plan and billing details.”, and rate limiting maps to rate_limit_exceeded. How to tell them apart, how to back off and how to switch models are below.

Updated 2026-10-08

Hit your subscription limit or keep getting 429s? Switch to an API billed per token and keep working with one key.

#Capacity signal#Not weekly quota#Not a 429#Switch model

Four key points

capacity

Error class

Literally “Selected model is at capacity. Please try a different model.”; in Codex's source as of 2026-10-07 it maps to the server error code server_is_overloaded — the model's load right now.

2026-09-17

Issue opened

user-filed issue #46189 in the official Codex repo records it; unresolved that day.

not quota

Easiest mistake

In the source, exhausted quota reads “Quota exceeded. Check your plan and billing details.” and rate limiting starts with “rate limit exceeded” — each has its own wording (checked 2026-10-07).

switch model

Action that works

Move to a sibling tier, or back off with jitter.

What this is

Codex returns this line when a model is saturated. It describes server-side capacity at that moment — not your quota state, not HTTP 429 throttling. The three have different wording and different remedies: quota means check the usage page, throttling means honour Retry-After, capacity means switch model or shift your window.

What happened

The user-filed Issue #46189 in the official Codex repository was opened 2026-09-17 with the literal error "Selected model is at capacity. Please try a different model.", and reports the account oscillating between this failure and a second one. Added 2026-10-07: the source of OpenAI's official Codex repository (release tag rust-v0.160.1, published 2026-10-05) defines this message as fixed text for the server error code server_is_overloaded; in the same file, exhausted quota and rate limiting are two other errors with different wording. In that version's source this error is classed as terminal and is not retried automatically, so switching the model or trying later is up to you.

Timeline

2026-09-17

2026-09-17 the user-filed issue #46189 opened, recording the error and the two alternating failure shapes.

2026-09-18

2026-09-18 re-check: the issue is still open, so this page states the issue text as-is.

2026-10-14

2026-10-14 is an unrelated but easily confused event: GPT-5.5 retires from ChatGPT / Work / Codex (API excluded).

Confirmed vs caution

Officially confirmed

The issue body says the account oscillates between "Selected model is at capacity" and a successful request whose model behaviour is abnormal, unresolved as of 2026-09-17. Added 2026-10-07 (Codex open-source code and official docs): ① in Codex source rust-v0.160.1 this text maps to the error code server_is_overloaded, a separate error from quota (Quota exceeded…) and rate limiting (rate_limit_exceeded); ② Codex's official model page recommends gpt-6.1-sol for complex coding and gpt-6-luna for focused, repeatable tasks; ③ since Codex CLI 0.159.1 (2026-09-29) the bundled default model is gpt-6.1-sol.

⚠️ Caution

OpenAI documented no capacity algorithm in the issue and gave no fix ETA; it was still unresolved on the day it was opened, 2026-09-17. Claims like "it always recovers in five minutes" or "some regions get priority" are not in the source.

Same vs differs

Same

All three make requests fail and all three should be treated as retryable in some way.

Differs

Capacity: switch model or shift time. Exhausted weekly quota: wait for reset or change plan — the wording points at usage. 429: rate limit, honour Retry-After. The wording separates them; don't merge them.

What to do

1) Read the error wording before blaming quota; 2) switch to a sibling tier (e.g. gpt-5.6-sol ↔ gpt-6-astra), or add exponential backoff with jitter; 3) if it persists, compare against the user-filed issue #46189 instead of guessing at config changes. 4) As of 2026-10-07, pick the fallback by the official guidance: Codex's model page recommends gpt-6.1-sol for complex coding and agentic workflows and gpt-6-luna for focused, repeatable tasks; since Codex CLI 0.159.1 (2026-09-29) the bundled default model is gpt-6.1-sol. Before switching, check status.openai.com: when it was fetched on 2026-10-07 the “Codex in ChatGPT Desktop” component was operational with no ongoing incidents — if the status page is clean and only one model returns this message, treat it as a capacity notice.

On QCode

As of 2026-10-07, QCode's docs connect Codex by setting base_url to https://api.qcode.cc/openai and wire_api to responses in ~/.codex/config.toml. GPT models all use this one endpoint: gpt-6.1-sol, gpt-6-sol, gpt-6-astra and gpt-5.6-sol sit under the same key, so switching model needs no endpoint change — handy for scripting “on a capacity notice, switch tier”. Claude and Chinese models don't go through the Responses endpoint Codex uses. Check usage and balance on the console's usage page; that is where you can confirm a quota problem. Billing is per token; per-model prices are on /models. QCode does not currently support gpt-6-luna or gpt-5.6-luna (both removed from /models); please use another GPT model, such as gpt-6.1-sol, gpt-6-sol or gpt-5.6-terra.

FAQ

Is this my weekly quota?

No. Quota exhaustion words itself differently and points at usage or plan. "Selected model is at capacity. Please try a different model" is about server-side capacity for that model right now.

Is it a 429?

No. Rate limiting has its own wording and retry semantics. A capacity signal is answered by switching model or shifting your window, not by waiting on a fixed header.

What helps fastest?

Switch to a sibling model and add exponential backoff with jitter. If several models share one key, the switch is a single field change. As of 2026-10-07, Codex's official model page recommends gpt-6.1-sol for complex coding and gpt-6-luna for focused, repeatable tasks. QCode does not currently support gpt-6-luna or gpt-5.6-luna (both removed from /models); please use another GPT model, such as gpt-6.1-sol, gpt-6-sol or gpt-5.6-terra.

Why does it also go weird-but-working?

The issue reports that alternating behaviour; OpenAI gives no cause in the thread, so this page draws no conclusion. Log both shapes with model id and timestamp and look for a pattern.

Same thing as the 2026-10-14 retirement?

No. The retirement removes a model from ChatGPT / ChatGPT Work / Codex and explicitly excludes the API; a capacity signal is a runtime load condition. Different wording, different remedy.

Should I open a ticket?

Reference the user-filed issue #46189 first. If you see the alternating failures too, attach timestamps, model id and the exact error text — far more useful than "it doesn't work".

Sources

user-filed issue #46189 in the official OpenAI Codex repo (2026-09-17; not an official statement) and the official release notes, captured 2026-09-18. Error text per the issue body. Sources added on 2026-10-07: the source of OpenAI's official Codex repository (release tag rust-v0.160.1, published 2026-10-05), the official Codex models page and changelog (learn.chatgpt.com), OpenAI's official status page (status.openai.com) and QCode's Codex setup page (docs.qcode.cc), all fetched 2026-10-07.

Make the tier switch one line

One QCode key covers gpt-6.1-sol, gpt-6-sol, gpt-6-astra and more — change the model in Codex to sidestep the peak.

Related

Only the error text present in the user-filed issue in OpenAI's official Codex repo and in the official release notes are reproduced; no inference about internal capacity policy. From 2026-10-07 the page also cites the source of OpenAI's official Codex repository and official docs, checked on 2026-10-07; official pages take precedence. Model availability is as listed on /models.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.