Troubleshooting · Availability

What 529 Overloaded actually means

As of 2026-10-08, Claude Code's official error reference says a 529 means the API is at capacity, usually only temporarily — it is not your usage limit and doesn't count against your quota. Changing your code will not help; changing your retry strategy will.

Updated 2026-10-08

Hit your subscription limit or keep getting 429s? Switch to an API billed per token and keep working with one key.

#Claude API#529#Overloaded#Retry strategy

Four points

529

Upstream congestion

Returned when server-side capacity is tight. Unrelated to your key’s quota, your parameters, or your request body.

429

You are being rate limited

This is the one that means you exceeded your own rate or quota. 529 and 429 need different handling — do not share a code path.

Backoff

The only effective client-side move

Exponential backoff with jitter. Retrying immediately worsens the congestion and gets you rate limited sooner.

Multi-route

Structural mitigation

When the same work can land on a different upstream, one congested provider stops meaning downtime. The cost is accepting model differences.

How 529 differs from 429 and 503

529 means the server is overloaded right now — a capacity problem, usually temporary. 429 means you hit your own rate or quota ceiling — a quota problem. 503 generally means the service is unavailable or under maintenance. None of them means your request was malformed, but the handling differs: 529 calls for backoff and retry, 429 calls for slowing down or raising limits, 503 calls for waiting and watching the status page. Collapsing all three into one catch block is the most common mistake in error handling.

Why discussion picks up periodically

Whenever a new model ships or a large migration happens, upstream capacity tightens for a while and 529 discussion rises with it. These waves usually ease as capacity is added — but for someone on a deadline, waiting for a vendor to expand is not an available answer.

Diagnostic order

First

Confirm the status code really is 529 and not 429. The response bodies differ; count them separately in your logs.

Second

Check the vendor status page. If it is a broad incident, no client-side change is more than mitigation.

Third

Check whether your own retry logic is amplifying the congestion: no jitter, no ceiling, and immediate retry all amplify. On Claude Code, since version 2.1.292 (2026-10-06) you can lengthen the base delay of 529 retries with CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS.

Before retrying, work out which code you actually got

What is confirmed

A 529 signals server-side capacity congestion and is unrelated to request content — this semantics is stated in vendor documentation. Claude Code's official error reference (checked 2026-10-08) shows the message “API Error: Repeated 529 Overloaded errors. The API is at capacity — this is usually temporary.” and states: “A 529 is not your usage limit and doesn't count against your quota.” It adds that capacity is tracked per model, so you can run /model and switch to a different model to keep working. The Claude Code CHANGELOG records that from 2.1.286 a failing model call sends at most 14 requests with the default retry settings, and that 2.1.292 added CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS. Exponential backoff with jitter is a widely validated client-side response.

Do not treat as fact

Claims that "529 is the vendor quietly throttling heavy users" have no supporting evidence and are not asserted here. The distinction between 529 and 429 is documented publicly, and conflating them leads you to the wrong fix.

Two responses

Client-side backoff only

Cheap to implement and measurably reduces failures. But while the upstream stays congested, all you can do is wait.

Let work land on multiple upstreams

One congested provider stops meaning downtime. The cost is accepting different model habits, and every route has its own failure modes.

Writing retries correctly

Three points. First, exponential backoff rather than a fixed interval: wait a second, then double each time. Second, add random jitter — if every client backs off by the same amount they retry in a synchronised wave, which prolongs the congestion. Third, set ceilings on both retry count and total wait, or a single congestion event will let your task queue grow without bound. Separately, when a streaming request breaks mid-way, do not blindly retry from the start; first check whether what you already received can be continued. If you use Claude Code: its documentation says it retries transient failures up to 10 times with exponential backoff before showing an error (rechecked 2026-10-08); version 2.1.286 (GitHub release 2026-09-30) changed how retries are counted — “one limit now covers a whole model call, so with the default retry settings a failing call sends at most 14 requests”; and version 2.1.292 (GitHub release 2026-10-06) added the environment variable CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS — in the changelog's words, “to set a longer base delay for the backoff when retrying an overloaded (529) request”. The changelog gives no default value for this variable, and this page does not guess one.

What QCode can and cannot do here

What it can do: one key lets you switch to another model family during congestion instead of waiting for a single model to recover. What it cannot do: we are not the model vendor and cannot change the vendor's server capacity, and requests through QCode can fail too. What we offer is the option to switch model families, not a promise of flawless service.

FAQ

Should I change my request parameters when I get a 529?

No. 529 is unrelated to request content; changing max_tokens, model parameters, or trimming the prompt will not make it go away. Change the retry strategy instead.

Can I use the same retry logic for 529 and 429?

Not advisable. 529 warrants backing off and retrying the same request; 429 means you need to lower your send rate or raise your limits, and blind retries keep tripping the limiter.

How long should I back off?

A common pattern is one second, doubling each attempt, with random jitter and ceilings on both retry count and total wait. The exact numbers depend on how much latency your task tolerates. In Claude Code (official docs and CHANGELOG as of 2026-10-08): CLAUDE_CODE_MAX_RETRIES sets the number of retry attempts, default 10; since 2.1.286 a failing model call sends at most 14 requests with the default retry settings; since 2.1.292 CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS lets you set a longer base delay for 529 retries.

What if a streaming request hits 529 mid-way?

First check whether what you already received is usable. Continue if you can; retry the whole thing only if you cannot. Blindly restarting duplicates billing and adds to the congestion.

Does QCode get 529s?

It can. Congestion on the vendor side is real and we are not immune to it. What we can do is let the same work move to another model family — not promise that failures will not happen.

Does multi-route mean no single point of failure?

No. Multiple routes lower the probability that one outage stops your work, but each route has its own failure modes and the routing layer itself can fail. It is mitigation, not elimination.

Sources

Status-code semantics follow each vendor’s official API documentation. Claude Code's retry behaviour and CLAUDE_CODE_OVERLOADED_RETRY_BASE_DELAY_MS come from Anthropic's Claude Code CHANGELOG (the 2.1.286 and 2.1.292 entries), the GitHub release pages of both versions and the Claude Code errors documentation, all fetched and checked on 2026-10-08. This page deliberately cites no third-party availability percentages — such figures move over time and we have no authoritative independent measurement to point to.

When one route is congested, take another

One key across multiple model families — switch and keep going when a route is under pressure.

Related reading

Status-code semantics here follow each vendor’s official documentation and may change between versions; the Claude Code details were checked on 2026-10-08 against the official CHANGELOG and documentation, which take precedence. Nothing on this page is an availability commitment.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.