> ## Documentation Index
> Fetch the complete documentation index at: https://docs.laozhang.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# How many concurrent requests can I send?

> LaoZhang API doesn't currently limit concurrency per account or token. Learn why you can still get a 429 and how to pace requests with a queue and backoff.

LaoZhang API doesn't currently limit concurrency per account or token, for text, image, or video models. Any change will be posted in [announcements](/en/changelog). Under heavy load you can still get a `429`, which comes from an upstream model's rate or capacity limit, so pace requests in your own code.

| Item | Details |
| - | - |
| Platform concurrency limit | None at present |
| Source of `429` | The upstream model's rate or capacity limit |
| Billing | Requests rejected with `429` before generation usually aren't billed |
| What to do | Cap concurrency with a local queue and back off on `429` |

## Handle a 429

<Steps>
  <Step title="Read the error body">
    The body of a `429` says whether it's a rate limit, a capacity shortage, or a quota problem. If it contains `insufficient_user_quota`, your balance can't cover the pre-authorized charge; see [Why a request can fail with remaining balance](/en/faq/balance-insufficient). Retrying won't help.
  </Step>

  <Step title="Back off and retry">
    Wait longer before each retry, for example 1, 2, 4, then 8 seconds, add a little random jitter, and cap the number of attempts. Don't resend the same request the moment a `429` arrives.
  </Step>

  <Step title="Send fewer requests at once">
    Limit in-flight requests with a local queue or semaphore. If `429` keeps coming, halve your concurrency and raise it again gradually.
  </Step>
</Steps>

## Cap concurrency with a semaphore

This Python example keeps at most 10 requests in flight and retries a `429` with exponential backoff:

```python theme={null}
import asyncio
import os
import random
from openai import AsyncOpenAI, RateLimitError

client = AsyncOpenAI(
    api_key=os.environ["LAOZHANG_API_KEY"],
    base_url="https://api.laozhang.ai/v1",
    max_retries=0,
)
limit = asyncio.Semaphore(10)  # requests in flight


async def ask(prompt, attempts=5):
    async with limit:
        for attempt in range(attempts):
            try:
                response = await client.chat.completions.create(
                    model="gpt-5.4-mini",
                    messages=[{"role": "user", "content": prompt}],
                )
                return response.choices[0].message.content
            except RateLimitError:
                if attempt == attempts - 1:
                    raise
                await asyncio.sleep(2 ** attempt + random.random())


async def main():
    prompts = [f"Describe the number {i} in one sentence" for i in range(50)]
    results = await asyncio.gather(*(ask(p) for p in prompts))
    print(results[:3])


asyncio.run(main())
```

Install the SDK with `pip install openai` and set `LAOZHANG_API_KEY` first. Image and reasoning requests take longer, so use lower concurrency for them and set timeouts as described in [API timeouts](/en/faq/request-timeout).

## Before a large batch

* Run a small batch first and check duration, success rate, and charges in [call logs](https://api.laozhang.ai/log).
* Video models run as asynchronous tasks. Poll at the interval each video model's guide suggests rather than as fast as you can.
* If you plan to send a large burst of requests, email the models, expected volume, and timing to [hi@laozhang.ai](mailto:hi@laozhang.ai) in advance so upstream capacity can be checked.

## FAQ

### Do more tokens give me more concurrency?

No. The platform doesn't limit concurrency per token, so splitting traffic across tokens doesn't raise any ceiling. Separate tokens per service or environment are for management and security; see [API key management](/en/faq/token-management).

### If there's no limit, why do requests slow down?

Upstream models queue requests at peak times, so each request takes longer, and higher concurrency means more requests waiting at once. For slow requests, first make sure your timeout is long enough, then consider lowering concurrency.

## Related pages

* [API timeouts](/en/faq/request-timeout)
* [Why a request can fail with remaining balance](/en/faq/balance-insufficient)
* [Check model availability and access](/en/faq/model-availability)
* [Connect your application to LaoZhang API: handle errors](/en/api-manual)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.