Skip to main content
Most timeouts happen because your client or a proxy in between stops waiting while the model is still generating. Closing the connection doesn’t cancel a request that has already reached the upstream provider, so it’s billed when generation finishes. Set a timeout that’s long enough the first time instead of a short one plus retries.
Retrying right after a timeout can cost you twice and still return nothing.
  • The request you dropped usually finishes and is billed.
  • Built-in SDK retries send the same request again after a timeout.
  • Before you retry, check the earlier request’s status and charge in call logs.

Suggested timeouts

Reasoning models think before they answer, and one request can run for several minutes. Examples:
  • gpt-6-sol and gpt-5.6-sol
  • gpt-5.5-pro and o3-pro
  • gemini-3.1-pro-preview

Stream long outputs

A non-streaming request returns only after the whole answer is generated, so your read timeout races the full generation time. With streaming (stream=True), output arrives as it’s generated:
  • The first data arrives quickly, and events keep coming after that.
  • The read timeout only needs to cover the gap between events; 90–120 seconds is usually enough.
  • Total time stays the same, but the connection no longer drops while it waits for the full answer.
If you stay with non-streaming requests, set the timeout to 300–600 seconds and expect longer generations to fail more often.

Turn off automatic retries for long requests

The official OpenAI Python SDK retries twice by default after a timeout and some other errors. For image and reasoning requests, set max_retries to 0 and let your own code decide when to retry:

Still timing out after raising the timeout

Work through these in order:
1

Confirm the new timeout is in effect

Some frameworks wrap the HTTP client in a timeout of their own. Print the effective settings and make sure you changed the one that’s used.
2

Check every hop

Any hop with a shorter timeout than the generation closes the connection first. Raise each one:
  • a reverse proxy you run, such as Nginx, whose proxy_read_timeout defaults to 60 seconds;
  • your cloud load balancer’s idle timeout;
  • your serverless function’s maximum run time;
  • your queue worker’s per-job timeout.
3

Tell timeouts apart from rate limits

A dropped connection or read timeout means you didn’t wait long enough. A 429 is a rate or capacity limit and has nothing to do with duration. Retry a 429 a few times with backoff, and contact support if it keeps happening.
4

Check the real duration in call logs

Open call logs, find the request’s duration and charge, and set your timeout above it with some headroom.

FAQ

Is a request billed if my client times out?

Yes, if the upstream provider finished generating. The disconnect happens on your side; the gateway and the provider keep going. Requests that return 429 or 503 before generation starts usually aren’t billed. Call logs show the actual charge, and refunds and balance adjustments follow the Terms of Service.

Can I fetch an image by ID after the connection drops?

No. Image endpoints are synchronous, and LaoZhang API doesn’t store generated results by default, so a dropped result is lost. If you need asynchronous behavior, put a task queue in your own backend, as described in Do image APIs return a task ID?

Does a very long timeout have side effects?

It doesn’t change billing, which depends only on usage or the number of calls, not on how long you wait. Long connections do hold a worker or pool slot, so under heavy load, run image and reasoning requests in a separate job queue.