Skip to main content
LaoZhang API bills image models per call or per returned image. Whether a request is charged depends on whether it reached generation and what the upstream returned, not on the HTTP status code alone. This page puts the rules for all three families in one table, last verified on October 10, 2026.

Charges by request outcome

“Not charged” follows each model guide: the Nano Banana guide says “usually not charged” and the GPT Image guide says “generally not charged”. The call log is authoritative; a request with no call-log entry produced no charge. Nano Banana’s HTTP 200 with no image is the only “no image” case among the three families that is charged: Google answers a moderation block with a short refusal instead of an error code, and LaoZhang API counts the upstream HTTP 200 as one call. See Avoid paying for empty responses for the usual triggers and how to avoid them.

Timeouts and retries

A client timeout is not a failure: the image may already have been generated and charged on the server. For batch jobs, handle requests in this order:
  1. Record the submit time, model ID, parameters, and response of every request.
  2. After a timeout or disconnect, find the request in the call log by time and model and check its status and charge.
  3. If the call log shows success, don’t resubmit; use the returned image or ask support to retrieve it.
  4. Retry 429 and 5xx with exponential backoff; each retry is a new request and is billed on its own if it succeeds.
  5. Don’t resend a Nano Banana HTTP 200 with no image unchanged. Change the prompt or references first; the same content is usually blocked again and charged every time.
  6. Higher resolutions and quality tiers take longer, so give the client at least 120 to 180 seconds.

Concurrency and rate limits

LaoZhang API doesn’t limit concurrency per account or token, and splitting traffic across tokens doesn’t raise any limit. When you get a 429 under load:
  • an error body containing insufficient_user_quota means your balance is too low, so top up first;
  • any other 429 comes from the upstream model’s rate or capacity limit, so halve your concurrency in a local queue and raise it gradually;
  • a persistent 503 means upstream capacity is short, not that your parameters are wrong, so wait and retry.
Concurrency and 429 has queue and semaphore examples.

Questioning a charge

Send support the following, and never a full API key:
  • your account email;
  • the request time, model ID, and the request identifier from the call log;
  • the response status and the redacted error;
  • the resolution you expect.
Contact hi@laozhang.ai or @laozhang_cn on Telegram. Refunds and balance adjustments follow section 6 of the Terms.

FAQ

Is a failed request always free?

Requests that return an error before generation aren’t charged: invalid parameters, authentication failures, rate limits, and upstream 5xx errors. The exception is a Nano Banana moderation block, where Google returns HTTP 200 instead of an error code, and that request is charged once.

How do I avoid paying twice in a batch job?

Give every target image a job ID, check the call log after a timeout, and retry only when the original request didn’t succeed. Never let a client timeout trigger an automatic resubmit. Grok Imagine can return up to 10 images in one request and bills each returned image.

How do I see what one request cost?

The call log records model, group, status, usage, and amount per request. Per-call models show a fixed amount, GPT Image official forwarding shows input and output tokens with the converted amount, and Nano Banana shows the search count and fee when Google Search grounding is on. How to read call logs explains the fields.