Codex error "Selected model is at capacity. Please try a different model": it is not your quota, and it is not a 429
Capacity, quota and 429 are three different things
As of 2026-10-07, this is Codex's capacity message — not exhausted quota and not a 429. In the source of OpenAI's official Codex repository (rust-v0.160.1, released 2026-10-05), “Selected model is at capacity. Please try a different model.” maps to the server error code server_is_overloaded; exhausted quota is a different line, “Quota exceeded. Check your plan and billing details.”, and rate limiting maps to rate_limit_exceeded. How to tell them apart, how to back off and how to switch models are below.
Updated 2026-10-08
Hit your subscription limit or keep getting 429s? Switch to an API billed per token and keep working with one key.
- Billed per token — live rates on /models
- Pay by card (Visa / Mastercard / AMEX), Apple Pay, Google Pay or crypto
- Self-service: top up and activate instantly
Four key points
Error class
Literally “Selected model is at capacity. Please try a different model.”; in Codex's source as of 2026-10-07 it maps to the server error code server_is_overloaded — the model's load right now.
Issue opened
user-filed issue #46189 in the official Codex repo records it; unresolved that day.
Easiest mistake
In the source, exhausted quota reads “Quota exceeded. Check your plan and billing details.” and rate limiting starts with “rate limit exceeded” — each has its own wording (checked 2026-10-07).
Action that works
Move to a sibling tier, or back off with jitter.
What this is
Codex returns this line when a model is saturated. It describes server-side capacity at that moment — not your quota state, not HTTP 429 throttling. The three have different wording and different remedies: quota means check the usage page, throttling means honour Retry-After, capacity means switch model or shift your window.
What happened
The user-filed Issue #46189 in the official Codex repository was opened 2026-09-17 with the literal error "Selected model is at capacity. Please try a different model.", and reports the account oscillating between this failure and a second one. Added 2026-10-07: the source of OpenAI's official Codex repository (release tag rust-v0.160.1, published 2026-10-05) defines this message as fixed text for the server error code server_is_overloaded; in the same file, exhausted quota and rate limiting are two other errors with different wording. In that version's source this error is classed as terminal and is not retried automatically, so switching the model or trying later is up to you.
Timeline
2026-09-17 the user-filed issue #46189 opened, recording the error and the two alternating failure shapes.
2026-09-18 re-check: the issue is still open, so this page states the issue text as-is.
2026-10-14 is an unrelated but easily confused event: GPT-5.5 retires from ChatGPT / Work / Codex (API excluded).
Confirmed vs caution
Officially confirmed
The issue body says the account oscillates between "Selected model is at capacity" and a successful request whose model behaviour is abnormal, unresolved as of 2026-09-17. Added 2026-10-07 (Codex open-source code and official docs): ① in Codex source rust-v0.160.1 this text maps to the error code server_is_overloaded, a separate error from quota (Quota exceeded…) and rate limiting (rate_limit_exceeded); ② Codex's official model page recommends gpt-6.1-sol for complex coding and gpt-6-luna for focused, repeatable tasks; ③ since Codex CLI 0.159.1 (2026-09-29) the bundled default model is gpt-6.1-sol.
⚠️ Caution
OpenAI documented no capacity algorithm in the issue and gave no fix ETA; it was still unresolved on the day it was opened, 2026-09-17. Claims like "it always recovers in five minutes" or "some regions get priority" are not in the source.
Same vs differs
Same
All three make requests fail and all three should be treated as retryable in some way.
Differs
Capacity: switch model or shift time. Exhausted weekly quota: wait for reset or change plan — the wording points at usage. 429: rate limit, honour Retry-After. The wording separates them; don't merge them.
What to do
1) Read the error wording before blaming quota; 2) switch to a sibling tier (e.g. gpt-5.6-sol ↔ gpt-6-astra), or add exponential backoff with jitter; 3) if it persists, compare against the user-filed issue #46189 instead of guessing at config changes. 4) As of 2026-10-07, pick the fallback by the official guidance: Codex's model page recommends gpt-6.1-sol for complex coding and agentic workflows and gpt-6-luna for focused, repeatable tasks; since Codex CLI 0.159.1 (2026-09-29) the bundled default model is gpt-6.1-sol. Before switching, check status.openai.com: when it was fetched on 2026-10-07 the “Codex in ChatGPT Desktop” component was operational with no ongoing incidents — if the status page is clean and only one model returns this message, treat it as a capacity notice.
On QCode
As of 2026-10-07, QCode's docs connect Codex by setting base_url to https://api.qcode.cc/openai and wire_api to responses in ~/.codex/config.toml. GPT models all use this one endpoint: gpt-6.1-sol, gpt-6-sol, gpt-6-astra and gpt-5.6-sol sit under the same key, so switching model needs no endpoint change — handy for scripting “on a capacity notice, switch tier”. Claude and Chinese models don't go through the Responses endpoint Codex uses. Check usage and balance on the console's usage page; that is where you can confirm a quota problem. Billing is per token; per-model prices are on /models. QCode does not currently support gpt-6-luna or gpt-5.6-luna (both removed from /models); please use another GPT model, such as gpt-6.1-sol, gpt-6-sol or gpt-5.6-terra.
FAQ
Is this my weekly quota?
No. Quota exhaustion words itself differently and points at usage or plan. "Selected model is at capacity. Please try a different model" is about server-side capacity for that model right now.
Is it a 429?
No. Rate limiting has its own wording and retry semantics. A capacity signal is answered by switching model or shifting your window, not by waiting on a fixed header.
What helps fastest?
Switch to a sibling model and add exponential backoff with jitter. If several models share one key, the switch is a single field change. As of 2026-10-07, Codex's official model page recommends gpt-6.1-sol for complex coding and gpt-6-luna for focused, repeatable tasks. QCode does not currently support gpt-6-luna or gpt-5.6-luna (both removed from /models); please use another GPT model, such as gpt-6.1-sol, gpt-6-sol or gpt-5.6-terra.
Why does it also go weird-but-working?
The issue reports that alternating behaviour; OpenAI gives no cause in the thread, so this page draws no conclusion. Log both shapes with model id and timestamp and look for a pattern.
Same thing as the 2026-10-14 retirement?
No. The retirement removes a model from ChatGPT / ChatGPT Work / Codex and explicitly excludes the API; a capacity signal is a runtime load condition. Different wording, different remedy.
Should I open a ticket?
Reference the user-filed issue #46189 first. If you see the alternating failures too, attach timestamps, model id and the exact error text — far more useful than "it doesn't work".
Sources
user-filed issue #46189 in the official OpenAI Codex repo (2026-09-17; not an official statement) and the official release notes, captured 2026-09-18. Error text per the issue body. Sources added on 2026-10-07: the source of OpenAI's official Codex repository (release tag rust-v0.160.1, published 2026-10-05), the official Codex models page and changelog (learn.chatgpt.com), OpenAI's official status page (status.openai.com) and QCode's Codex setup page (docs.qcode.cc), all fetched 2026-10-07.
Make the tier switch one line
One QCode key covers gpt-6.1-sol, gpt-6-sol, gpt-6-astra and more — change the model in Codex to sidestep the peak.
Related
GPT-5.5 retirement
A 2026-10-14 model removal, not a capacity signal.
Codex disable_response_storage: removed
Added for ZDR orgs in 2025-04, removed by PR #3212 on 2025-09-05; not in the current config-reference, and old configs get a warning.
Codex model_catalog_json configuration guide
Loaded at startup, replaces the bundled catalog, profile wins; the 0.160.0 explicit provider catalog change and /model troubleshooting.
Only the error text present in the user-filed issue in OpenAI's official Codex repo and in the official release notes are reproduced; no inference about internal capacity policy. From 2026-10-07 the page also cites the source of OpenAI's official Codex repository and official docs, checked on 2026-10-07; official pages take precedence. Model availability is as listed on /models.