> ## Documentation Index
> Fetch the complete documentation index at: https://docs.laozhang.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# API timeouts: how long should the timeout be?

> Suggested timeouts for text, reasoning, and image requests on LaoZhang API, why a dropped request is still billed, and how to troubleshoot.

Most timeouts happen because your client or a proxy in between stops waiting while the model is still generating. Closing the connection doesn't cancel a request that has already reached the upstream provider, so it's billed when generation finishes. Set a timeout that's long enough the first time instead of a short one plus retries.

<Warning>
  **Retrying right after a timeout can cost you twice and still return nothing.**

  * The request you dropped usually finishes and is billed.
  * Built-in SDK retries send the same request again after a timeout.
  * Before you retry, check the earlier request's status and charge in [call logs](https://api.laozhang.ai/log).
</Warning>

## Suggested timeouts

| Workload | Suggested timeout | Notes |
| - | - | - |
| Regular chat | 60–120 s | Usually returns within seconds |
| Reasoning models, long output | Stream it; 300–600 s without streaming | Higher reasoning effort takes longer |
| Image generation and editing | 360 s | Image endpoints are synchronous, so a dropped connection loses the result |
| 4K output, many reference images | 600 s | Slower at peak times |
| Video generation | Poll the task | Submit and status calls return quickly; see each video model's guide |

Reasoning models think before they answer, and one request can run for several minutes. Examples:

* `gpt-6-sol` and `gpt-5.6-sol`
* `gpt-5.5-pro` and `o3-pro`
* `gemini-3.1-pro-preview`

## Stream long outputs

A non-streaming request returns only after the whole answer is generated, so your read timeout races the full generation time. With streaming (`stream=True`), output arrives as it's generated:

* The first data arrives quickly, and events keep coming after that.
* The read timeout only needs to cover the gap between events; 90–120 seconds is usually enough.
* Total time stays the same, but the connection no longer drops while it waits for the full answer.

If you stay with non-streaming requests, set the timeout to 300–600 seconds and expect longer generations to fail more often.

## Turn off automatic retries for long requests

The official OpenAI Python SDK retries twice by default after a timeout and some other errors. For image and reasoning requests, set `max_retries` to 0 and let your own code decide when to retry:

<Tabs>
  <Tab title="Python">
    ```python theme={null}
    import os
    from openai import OpenAI

    client = OpenAI(
        api_key=os.environ["LAOZHANG_API_KEY"],
        base_url="https://api.laozhang.ai/v1",
        max_retries=0,  # no automatic retries, so a timeout isn't billed twice
    )

    response = client.chat.completions.create(
        model="gpt-6-sol",
        messages=[{"role": "user", "content": "Analyze the time complexity of this code"}],
        timeout=600,  # give reasoning models enough time
    )
    print(response.choices[0].message.content)
    ```
  </Tab>

  <Tab title="Node.js">
    ```javascript theme={null}
    import OpenAI from "openai";

    const client = new OpenAI({
      apiKey: process.env.LAOZHANG_API_KEY,
      baseURL: "https://api.laozhang.ai/v1",
      timeout: 600 * 1000, // milliseconds
      maxRetries: 0,
    });

    const response = await client.chat.completions.create({
      model: "gpt-6-sol",
      messages: [{ role: "user", content: "Write a 3,000-word technical analysis" }],
    });
    console.log(response.choices[0].message.content);
    ```
  </Tab>

  <Tab title="cURL">
    ```bash theme={null}
    # --max-time caps the whole request, in seconds
    curl https://api.laozhang.ai/v1/images/generations \
      -H "Authorization: Bearer $LAOZHANG_API_KEY" \
      -H "Content-Type: application/json" \
      --max-time 360 \
      -d '{"model": "gpt-image-2.5-flare-vip", "prompt": "A calm mountain lake at sunrise"}'
    ```
  </Tab>
</Tabs>

## Still timing out after raising the timeout

Work through these in order:

<Steps>
  <Step title="Confirm the new timeout is in effect">
    Some frameworks wrap the HTTP client in a timeout of their own. Print the effective settings and make sure you changed the one that's used.
  </Step>

  <Step title="Check every hop">
    Any hop with a shorter timeout than the generation closes the connection first. Raise each one:

    * a reverse proxy you run, such as Nginx, whose `proxy_read_timeout` defaults to 60 seconds;
    * your cloud load balancer's idle timeout;
    * your serverless function's maximum run time;
    * your queue worker's per-job timeout.
  </Step>

  <Step title="Tell timeouts apart from rate limits">
    A dropped connection or read timeout means you didn't wait long enough. A `429` is a rate or capacity limit and has nothing to do with duration. Retry a `429` a few times with backoff, and contact support if it keeps happening.
  </Step>

  <Step title="Check the real duration in call logs">
    Open [call logs](https://api.laozhang.ai/log), find the request's duration and charge, and set your timeout above it with some headroom.
  </Step>
</Steps>

## FAQ

### Is a request billed if my client times out?

Yes, if the upstream provider finished generating. The disconnect happens on your side; the gateway and the provider keep going. Requests that return `429` or `503` before generation starts usually aren't billed. Call logs show the actual charge, and refunds and balance adjustments follow the [Terms of Service](https://www.laozhang.ai/terms).

### Can I fetch an image by ID after the connection drops?

No. Image endpoints are synchronous, and LaoZhang API doesn't store generated results by default, so a dropped result is lost. If you need asynchronous behavior, put a task queue in your own backend, as described in [Do image APIs return a task ID?](/en/faq/image-api-sync)

### Does a very long timeout have side effects?

It doesn't change billing, which depends only on usage or the number of calls, not on how long you wait. Long connections do hold a worker or pool slot, so under heavy load, run image and reasoning requests in a separate job queue.

## Related pages

* [Do image APIs return a task ID?](/en/faq/image-api-sync)
* [Connect your application to LaoZhang API: handle errors](/en/api-manual)
* [How to review and use call logs](/en/faq/call-logs)
* [Text generation API: streaming in each protocol](/en/api-capabilities/text-generation)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.