Skip to main content
LaoZhang API doesn’t currently limit concurrency per account or token, for text, image, or video models. Any change will be posted in announcements. Under heavy load you can still get a 429, which comes from an upstream model’s rate or capacity limit, so pace requests in your own code.

Handle a 429

1

Read the error body

The body of a 429 says whether it’s a rate limit, a capacity shortage, or a quota problem. If it contains insufficient_user_quota, your balance can’t cover the pre-authorized charge; see Why a request can fail with remaining balance. Retrying won’t help.
2

Back off and retry

Wait longer before each retry, for example 1, 2, 4, then 8 seconds, add a little random jitter, and cap the number of attempts. Don’t resend the same request the moment a 429 arrives.
3

Send fewer requests at once

Limit in-flight requests with a local queue or semaphore. If 429 keeps coming, halve your concurrency and raise it again gradually.

Cap concurrency with a semaphore

This Python example keeps at most 10 requests in flight and retries a 429 with exponential backoff:
Install the SDK with pip install openai and set LAOZHANG_API_KEY first. Image and reasoning requests take longer, so use lower concurrency for them and set timeouts as described in API timeouts.

Before a large batch

  • Run a small batch first and check duration, success rate, and charges in call logs.
  • Video models run as asynchronous tasks. Poll at the interval each video model’s guide suggests rather than as fast as you can.
  • If you plan to send a large burst of requests, email the models, expected volume, and timing to hi@laozhang.ai in advance so upstream capacity can be checked.

FAQ

Do more tokens give me more concurrency?

No. The platform doesn’t limit concurrency per token, so splitting traffic across tokens doesn’t raise any ceiling. Separate tokens per service or environment are for management and security; see API key management.

If there’s no limit, why do requests slow down?

Upstream models queue requests at peak times, so each request takes longer, and higher concurrency means more requests waiting at once. For slow requests, first make sure your timeout is long enough, then consider lowering concurrency.