> ## Documentation Index
> Fetch the complete documentation index at: https://docs.laozhang.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Nano Banana 2.1 API: generate and edit images

> Generate and edit 1K to 4K images with gemini-nano-banana-2.1 on LaoZhang API at 0.045 USD per call: 14 aspect ratios, 14 references, and runnable code.

Nano Banana 2.1 (`gemini-nano-banana-2.1`) is Google's October 2026 update to Nano Banana 2, with sharper text rendering, more accurate infographic layouts, and better consistency across multi-turn edits. On LaoZhang API it costs \$0.045 per call, less than Nano Banana 2.

## What Nano Banana 2.1 supports

| Item | Details |
| - | - |
| Model ID | `gemini-nano-banana-2.1` (stable) |
| Price | \$0.045 per call (public pricing as of October 7, 2026; see [Models and pricing](/en/models)) |
| Billing | Every HTTP 200 response is charged once, even with no image; see [billing rules](/en/api-capabilities/nano-banana-image#billing) |
| Token | `default` group, billing mode Usage first (recommended) or Per-call |
| Endpoints | Gemini-native `generateContent` sets the aspect ratio and output size; OpenAI-compatible chat completions returns 1K images |
| Output sizes | `1K` (default), `2K`, `4K`; `512` isn't supported |
| Aspect ratios | 14: `1:1`, `1:4`, `1:8`, `2:3`, `3:2`, `3:4`, `4:1`, `4:3`, `4:5`, `5:4`, `8:1`, `9:16`, `16:9`, `21:9` |
| Reference images | Up to 14: at most 10 object images and 4 character images |
| Thinking | `minimal`, `medium` (default), or `high` |
| Google Search grounding | Web search and image search, \$0.014 per search |

To compare all five Nano Banana models, see the [Nano Banana overview](/en/api-capabilities/nano-banana-image).

## Nano Banana 2.1 or Nano Banana 2?

Both models take the same request and return the same response shape, so switching means changing only the model ID.

| | Nano Banana 2.1 | Nano Banana 2 |
| - | - | - |
| Model ID | `gemini-nano-banana-2.1` | `gemini-3.1-flash-image` |
| LaoZhang API price | \$0.045 per call | \$0.055 per call |
| Output sizes | `1K`, `2K`, `4K` | `512`, `1K`, `2K`, `4K` |
| Thinking levels | `minimal`, `medium` (default), `high` | `minimal` (default), `high` |
| Ratios, references, grounding | Same | Same |

* **New projects and 1K to 4K output**: use Nano Banana 2.1. It costs less per call and handles text and layout better.
* **512px previews**: stay on [Nano Banana 2](/en/api-capabilities/nano-banana2-image). Nano Banana 2.1 returns HTTP 400 for `512`.
* Nano Banana 2 remains available at the same price, so existing integrations don't have to move.

## Before you call

<Steps>
  <Step title="Create an API key">
    In [Token management](https://api.laozhang.ai/token), create a key in the `default` group and set **Billing mode** (计费模式) to **Usage first** (按量优先, recommended) or **Per-call** (按次计费). A Usage-based (按量计费) key can't call per-call models such as Nano Banana 2.1.
  </Step>

  <Step title="Set the key and install the dependencies">
    ```bash theme={null}
    export LAOZHANG_API_KEY="sk-..."  # replace with your LaoZhang API key
    python -m pip install openai requests
    ```
  </Step>

  <Step title="Add a save function for cURL">
    The cURL examples need `jq`. This shell function saves the last non-thought image in a Gemini-native response, with the extension that matches its `mimeType`:

    ```bash theme={null}
    save_image() {
      local image='[.candidates[]?.content.parts[]? | select(.inlineData and (.thought | not))][-1].inlineData'
      local ext
      case "$(jq -r "$image.mimeType" "$1")" in
        image/png) ext=png ;;
        image/jpeg) ext=jpg ;;
        image/webp) ext=webp ;;
        *) jq '{error, promptFeedback, finishReasons: [.candidates[]?.finishReason]}' "$1"; return 1 ;;
      esac
      jq -r "$image.data" "$1" | base64 --decode > "$2.$ext" && echo "Saved $2.$ext"
    }
    ```

    When the response holds no image, the function prints the error, `promptFeedback`, and `finishReason` fields instead.
  </Step>
</Steps>

A request usually takes 15 to 40 seconds, and 4K takes longer. The examples allow 300 seconds; give your client and any proxy the same headroom.

## Generate an image

<Tabs>
  <Tab title="cURL">
    ```bash theme={null}
    curl --fail-with-body --max-time 300 \
      "https://api.laozhang.ai/v1beta/models/gemini-nano-banana-2.1:generateContent" \
      -H "Authorization: Bearer $LAOZHANG_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "contents": [{
          "role": "user",
          "parts": [{"text": "Generate an image of a corner cafe after rain at dusk, warm light on the wet pavement, with a neon sign that reads \"BANANA CAFE\"."}]
        }],
        "generationConfig": {
          "responseModalities": ["IMAGE"],
          "imageConfig": {"aspectRatio": "16:9", "imageSize": "2K"}
        }
      }' \
      -o nb21-response.json

    save_image nb21-response.json nb21-cafe
    ```
  </Tab>

  <Tab title="Python">
    Save this script as `nano_banana_21.py`. It generates an image from a prompt, or edits the images you list after the prompt:

    ```python theme={null}
    """Generate or edit images with Nano Banana 2.1 (gemini-nano-banana-2.1)."""
    import argparse
    import base64
    import os
    import sys
    from pathlib import Path

    import requests

    URL = "https://api.laozhang.ai/v1beta/models/gemini-nano-banana-2.1:generateContent"
    RATIOS = ["1:1", "1:4", "1:8", "2:3", "3:2", "3:4", "4:1", "4:3", "4:5", "5:4", "8:1", "9:16", "16:9", "21:9"]
    SIZES = ["1K", "2K", "4K"]
    MIME_TYPES = {".png": "image/png", ".jpg": "image/jpeg", ".jpeg": "image/jpeg", ".webp": "image/webp"}
    EXTENSIONS = {"image/png": "png", "image/jpeg": "jpg", "image/webp": "webp"}

    parser = argparse.ArgumentParser(description="Generate an image, or edit the images you pass.")
    parser.add_argument("prompt")
    parser.add_argument("images", nargs="*", help="up to 14 PNG, JPEG, or WebP reference images")
    parser.add_argument("--ratio", choices=RATIOS, help="aspect ratio; omit to let the model choose")
    parser.add_argument("--size", default="1K", choices=SIZES)
    parser.add_argument("--thinking", choices=["minimal", "medium", "high"], help="omit for the default, medium")
    parser.add_argument("--out", default="nano-banana-21", help="output file name without extension")
    args = parser.parse_args()
    if len(args.images) > 14:
        sys.exit("Send at most 14 reference images.")

    parts = [{"text": args.prompt}]
    for path in args.images:
        mime = MIME_TYPES.get(Path(path).suffix.lower())
        if mime is None:
            sys.exit(f"Use a PNG, JPEG, or WebP file: {path}")
        data = base64.b64encode(Path(path).read_bytes()).decode("ascii")
        parts.append({"inlineData": {"mimeType": mime, "data": data}})

    config = {"responseModalities": ["IMAGE"], "imageConfig": {"imageSize": args.size}}
    if args.ratio:
        config["imageConfig"]["aspectRatio"] = args.ratio
    if args.thinking:
        config["thinkingConfig"] = {"thinkingLevel": args.thinking}

    response = requests.post(
        URL,
        headers={"Authorization": f"Bearer {os.environ['LAOZHANG_API_KEY']}"},
        json={"contents": [{"role": "user", "parts": parts}], "generationConfig": config},
        timeout=300,
    )
    if response.status_code != 200:
        sys.exit(f"HTTP {response.status_code}: {response.text[:500]}")

    result = response.json()
    images = [
        part["inlineData"]
        for candidate in result.get("candidates") or []
        for part in (candidate.get("content") or {}).get("parts") or []
        if "inlineData" in part and not part.get("thought")
    ]
    if not images:
        reasons = [candidate.get("finishReason") for candidate in result.get("candidates") or []]
        sys.exit(f"No image returned: {result.get('promptFeedback') or reasons}")
    for number, image in enumerate(images, start=1):
        suffix = "" if number == 1 else f"-{number}"
        output = Path(f"{args.out}{suffix}.{EXTENSIONS.get(image['mimeType'], 'bin')}")
        output.write_bytes(base64.b64decode(image["data"]))
        print("Saved", output)
    ```

    Generate a 4K, 16:9 image:

    ```bash theme={null}
    python nano_banana_21.py "Generate an image of a corner cafe after rain at dusk, warm light on the wet pavement, with a neon sign that reads \"BANANA CAFE\"." --ratio 16:9 --size 4K --out nb21-cafe
    ```
  </Tab>
</Tabs>

The image comes back as Base64 in `candidates[0].content.parts[].inlineData.data`, and `mimeType` gives its format. Both examples decode it and save a matching PNG, JPEG, or WebP file.

Without `aspectRatio`, the model picks the frame from the prompt. Common output sizes:

* 1:1: 1024×1024 at `1K`, 2048×2048 at `2K`, 4096×4096 at `4K`
* 16:9: 1376×768 at `1K`, 2752×1536 at `2K`
* Other ratios: see Google's [image generation guide](https://ai.google.dev/gemini-api/docs/image-generation)

## Edit images

Send the instruction first, then the reference images as `inlineData` parts, and refer to the images by their order in the prompt ("the first image", "the second image"). Nano Banana 2.1 accepts up to 14 references: at most 10 object images and 4 character images.

<Tabs>
  <Tab title="cURL">
    This request combines two local files. Build the body with `jq` so large images don't exceed shell argument limits, and keep each `mimeType` consistent with its file:

    ```bash theme={null}
    base64 < product.png | tr -d '\n' > product.b64
    base64 < scene.jpg | tr -d '\n' > scene.b64
    jq -n --rawfile product product.b64 --rawfile scene scene.b64 '{
      contents: [{
        role: "user",
        parts: [
          {text: "Place the product from the first image on the table in the second image. Match the scene lighting and keep the product label legible."},
          {inlineData: {mimeType: "image/png", data: $product}},
          {inlineData: {mimeType: "image/jpeg", data: $scene}}
        ]
      }],
      generationConfig: {
        responseModalities: ["IMAGE"],
        imageConfig: {aspectRatio: "3:2", imageSize: "2K"}
      }
    }' > nb21-edit-request.json

    curl --fail-with-body --max-time 300 \
      "https://api.laozhang.ai/v1beta/models/gemini-nano-banana-2.1:generateContent" \
      -H "Authorization: Bearer $LAOZHANG_API_KEY" \
      -H "Content-Type: application/json" \
      --data-binary @nb21-edit-request.json \
      -o nb21-edit-response.json

    save_image nb21-edit-response.json nb21-product-scene
    ```
  </Tab>

  <Tab title="Python">
    Use `nano_banana_21.py` from [Generate an image](#generate-an-image) and list the files after the prompt:

    ```bash theme={null}
    # Edit one image
    python nano_banana_21.py "Turn this scene into a snowy morning. Keep the building, the sign text, and the composition." nb21-cafe.jpg --ratio 16:9 --size 2K --out nb21-snow

    # Combine three references: product, scene, and a third image for the look
    python nano_banana_21.py "Place the product from the first image on the table in the second image, painted in the style of the third image." product.png scene.jpg style.jpg --ratio 3:2 --size 2K --out nb21-composite
    ```
  </Tab>
</Tabs>

Say what must stay unchanged, and describe one change per request when you need precise control. To keep people consistent across a series, pass their photos as character references (up to 4).

## Ground images in Google Search

Add the `googleSearch` tool when an image depends on current facts or on what a real landmark looks like. This request turns on both web search and image search, and allows `TEXT` so any text the model writes comes back too:

```bash theme={null}
curl --fail-with-body --max-time 300 \
  "https://api.laozhang.ai/v1beta/models/gemini-nano-banana-2.1:generateContent" \
  -H "Authorization: Bearer $LAOZHANG_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [{"role": "user", "parts": [{"text": "Generate a watercolor illustration of the tallest building in the world today, with its name written below it."}]}],
    "tools": [{"googleSearch": {"searchTypes": {"webSearch": {}, "imageSearch": {}}}}],
    "generationConfig": {
      "responseModalities": ["TEXT", "IMAGE"],
      "imageConfig": {"aspectRatio": "2:3"}
    }
  }' \
  -o nb21-grounded-response.json

save_image nb21-grounded-response.json nb21-tallest-building
jq '.candidates[0].groundingMetadata | {webSearchQueries, imageSearchQueries}' nb21-grounded-response.json
```

* `searchTypes` must be an object. Without it, the tool uses web search only; `imageSearch` alone uses image search only.
* The model decides whether to search and which type to use. The queries it ran appear in `groundingMetadata`.
* Image search can't currently use real-world images of people as references.
* If you show grounded results to end users, display search suggestions as Google's [grounding guide](https://ai.google.dev/gemini-api/docs/google-search) requires.
* Each search adds \$0.014 to the \$0.045 call, so a request that runs 2 searches costs \$0.073. Call logs show the number of searches; see the [Nano Banana billing rules](/en/api-capabilities/nano-banana-image#billing).

## Use the OpenAI-compatible route

If your app already uses the OpenAI SDK, point `base_url` at `https://api.laozhang.ai/v1` and set `model` to `gemini-nano-banana-2.1`. This route has no ratio or size settings: it returns a 1K image whose frame the model picks from the prompt, often 16:9. For a fixed ratio or 2K and 4K, use the Gemini-native route.

The image comes back as a Base64 data URL inside `message.content`. Save this script as `nano_banana_21_chat.py`:

```python theme={null}
"""Generate or edit an image with Nano Banana 2.1 over the OpenAI-compatible route."""
import argparse
import base64
import os
import re
import sys
from pathlib import Path

from openai import OpenAI

MIME_TYPES = {".png": "image/png", ".jpg": "image/jpeg", ".jpeg": "image/jpeg", ".webp": "image/webp"}
EXTENSIONS = {"png": "png", "jpeg": "jpg", "jpg": "jpg", "webp": "webp"}

parser = argparse.ArgumentParser(description="Generate an image, or edit the images you pass.")
parser.add_argument("prompt")
parser.add_argument("images", nargs="*", help="PNG, JPEG, or WebP images to edit")
parser.add_argument("--out", default="nano-banana-21-chat", help="output file name without extension")
args = parser.parse_args()

content = args.prompt
if args.images:  # edit: send the instruction and the local images as data URLs
    content = [{"type": "text", "text": args.prompt}]
    for path in args.images:
        mime = MIME_TYPES.get(Path(path).suffix.lower())
        if mime is None:
            sys.exit(f"Use a PNG, JPEG, or WebP file: {path}")
        encoded = base64.b64encode(Path(path).read_bytes()).decode("ascii")
        content.append({"type": "image_url", "image_url": {"url": f"data:{mime};base64,{encoded}"}})

# max_retries=0: an automatic retry after a timeout could generate and bill the image twice
client = OpenAI(api_key=os.environ["LAOZHANG_API_KEY"], base_url="https://api.laozhang.ai/v1", timeout=300, max_retries=0)
reply = client.chat.completions.create(
    model="gemini-nano-banana-2.1",
    messages=[{"role": "user", "content": content}],
)
text = reply.choices[0].message.content or ""

found = re.findall(r"data:image/(png|jpe?g|webp);base64,([A-Za-z0-9+/=]+)", text)
if not found:
    sys.exit(f"No image returned: {text[:300]}")
for number, (kind, encoded) in enumerate(found, start=1):
    suffix = "" if number == 1 else f"-{number}"
    output = Path(f"{args.out}{suffix}.{EXTENSIONS[kind]}")
    output.write_bytes(base64.b64decode(encoded + "=" * (-len(encoded) % 4)))
    print("Saved", output)
```

Generate an image, then edit a local file:

```bash theme={null}
python nano_banana_21_chat.py "Generate an image of a latte on a wooden table, with cat latte art, shot from above." --out nb21-chat-latte
python nano_banana_21_chat.py "Transform this photo into a Van Gogh style oil painting. Keep the composition." photo.jpg --out nb21-chat-painting
```

## Errors and billing

Invalid parameters return HTTP 400 and aren't charged:

| Request | Result |
| - | - |
| `imageSize` set to `512` | 400 with `Image size 512 is not supported for this model` |
| `aspectRatio` outside the 14 ratios, such as `7:3` | 400, possibly with an empty error message; check the ratio against the list above |

* Each call costs \$0.045, whatever the output size, thinking level, number of references, or whether it generates or edits.
* Google Search grounding adds \$0.014 per search.
* A request can return HTTP 200 without an image, for example when the prompt doesn't ask for one or a safety check blocks the output. It is charged as one call. Check `finishReason`, then change the prompt or references before resending; see [Avoid paying for empty responses](/en/api-capabilities/nano-banana-image#avoid-paying-for-empty-responses).
* Retry 429 and 5xx responses with backoff, and check [call logs](/en/faq/call-logs) before resending a request that timed out.

## FAQ

### What do I change to move from Nano Banana 2?

Change the model ID from `gemini-3.1-flash-image` to `gemini-nano-banana-2.1`; the request and response stay the same. Then check two things:

* Requests that set `"imageSize": "512"` need `1K` instead, or should stay on Nano Banana 2.
* Without a `thinkingLevel`, Nano Banana 2.1 thinks at `medium`, while Nano Banana 2 defaults to `minimal`. Set `minimal` explicitly if you want the faster behavior.

### How do I change the thinking level, and does it affect the price?

On the Gemini-native route, set `generationConfig.thinkingConfig.thinkingLevel` to `minimal`, `medium`, or `high`. Higher levels spend more effort on layout and text and take longer.

```json theme={null}
"thinkingConfig": {"thinkingLevel": "high"}
```

The price stays at \$0.045 per call for all three levels; Google itself bills thinking tokens on top. If you also set `"includeThoughts": true`, the response adds thought text and draft images marked `"thought": true`. The `save_image` function and the Python script skip them.

### Does 4K cost more than 1K?

No. Every call costs \$0.045. Google bills by output tokens, about \$0.113 for a 4K image and \$0.0336 for 1K, according to its [pricing page](https://ai.google.dev/gemini-api/docs/pricing). If you only need 1K images, [Nano Banana 2 Lite](/en/api-capabilities/nano-banana-2-lite-api) costs less per call.

### Why does the OpenAI-compatible route return landscape images?

That route can't pass an aspect ratio, so the model chooses the frame from the prompt, often 16:9. For square or portrait output, use the Gemini-native route and set `aspectRatio`.

## Related pages and sources

* [Nano Banana overview](/en/api-capabilities/nano-banana-image) — compare the five models, the two protocols, and billing
* [Nano Banana 2](/en/api-capabilities/nano-banana2-image) — when you need 512px output
* [Nano Banana 2 Lite](/en/api-capabilities/nano-banana-2-lite-api) — low-cost 1K images
* [Nano Banana Pro](/en/api-capabilities/nano-banana-pro-image) — complex layouts and brand-critical assets
* [Gemini protocol](/en/api-reference/gemini) — native request structure and SDK setup
* [Google: Gemini Nano Banana 2.1](https://ai.google.dev/gemini-api/docs/models/gemini-nano-banana-2.1) — official capabilities and limits
* [Google: Nano Banana image generation](https://ai.google.dev/gemini-api/docs/image-generation) — sizes, aspect ratios, and reference limits
* [Call logs](/en/faq/call-logs) — check a request's status and charge


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.