429, which comes from an upstream model’s rate or capacity limit, so pace requests in your own code.
Handle a 429
1
Read the error body
The body of a
429 says whether it’s a rate limit, a capacity shortage, or a quota problem. If it contains insufficient_user_quota, your balance can’t cover the pre-authorized charge; see Why a request can fail with remaining balance. Retrying won’t help.2
Back off and retry
Wait longer before each retry, for example 1, 2, 4, then 8 seconds, add a little random jitter, and cap the number of attempts. Don’t resend the same request the moment a
429 arrives.3
Send fewer requests at once
Limit in-flight requests with a local queue or semaphore. If
429 keeps coming, halve your concurrency and raise it again gradually.Cap concurrency with a semaphore
This Python example keeps at most 10 requests in flight and retries a429 with exponential backoff:
pip install openai and set LAOZHANG_API_KEY first. Image and reasoning requests take longer, so use lower concurrency for them and set timeouts as described in API timeouts.
Before a large batch
- Run a small batch first and check duration, success rate, and charges in call logs.
- Video models run as asynchronous tasks. Poll at the interval each video model’s guide suggests rather than as fast as you can.
- If you plan to send a large burst of requests, email the models, expected volume, and timing to hi@laozhang.ai in advance so upstream capacity can be checked.