Model Catalog and Pricing
This page lists 281 online LaoZhang API models with input, output, cache-read, per-call, and tiered prices, token groups, and supported endpoints. The data was updated on 2026-07-25, 22:52 (UTC+8). Confirm final availability and charges in the signed-in console and call logs.
Online models, cache-read prices, and tiered pricing rules are refreshed from current system configuration. Use the signed-in console to confirm account groups and current prices, and call logs to confirm actual charges.
Pricing notes
All amounts on this page are shown in US dollars. Usage-based input, output, and cache-read prices are per million tokens ($/1M). Per-call models are priced for each call. The table shows default list prices; account contracts or dedicated routes may use different prices.
Top-up bonus
A single top-up of US$700 or more receives a 5% balance bonus. The bonus does not rewrite the listed model price. Confirm the credited balance and actual charges in the console.
A model can support multiple token groups with different routes, permissions, or billing modes. Confirm the groups and actual prices available to your account in the console.
Field reference
| Field | Meaning |
|---|---|
| Model ID | The model name used in API requests. A † mark identifies a model with token-volume pricing tiers. |
| Input / output | The usage-based US dollar price per 1M tokens. |
| Cache read | The input price when a supported prompt cache is hit. “—” means no current cache-read price is listed. |
| Per-call price | The US dollar price per call. A call can mean a request, image, or task depending on the model documentation. |
| Token groups | Available group names. Account access and actual prices require a console check. |
| Supported endpoints | Declared protocol entry points. This does not guarantee that every client parameter is fully compatible. |
Online model catalog
Models within each vendor are ordered from newer to older versions. Use your browser find command or the search box above. Wide tables scroll horizontally on small screens.
OpenAI 93
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
gpt-5.6-luna † | Usage | $1 / 1M tokens | $6 / 1M tokens | $0.1 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.6-sol † | Usage | $5 / 1M tokens | $30 / 1M tokens | $0.5 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.6-terra † | Usage | $2.5 / 1M tokens | $15 / 1M tokens | $0.25 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.5 † | Usage | $5 / 1M tokens | $30 / 1M tokens | $0.5 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.5-pro † | Usage | $30 / 1M tokens | $180 / 1M tokens | $3 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.4 † | Usage | $2.5 / 1M tokens | $15 / 1M tokens | $0.25 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.4-mini | Usage | $0.75 / 1M tokens | $4.5 / 1M tokens | $0.075 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.4-nano | Usage | $0.2 / 1M tokens | $1.25 / 1M tokens | $0.02 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.4-pro † | Usage | $30 / 1M tokens | $180 / 1M tokens | $3 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.3 | Usage | $75 / 1M tokens | $450 / 1M tokens | $7.5 / 1M tokens | — | default | Chat Completions |
gpt-5.2 | Usage | $1.75 / 1M tokens | $14 / 1M tokens | $0.175 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.2-pro | Usage | $21 / 1M tokens | $168 / 1M tokens | $2.1 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.2-2025-12-11 | Usage | $1.75 / 1M tokens | $14 / 1M tokens | $0.175 / 1M tokens | — | default | Chat Completions |
gpt-5.1 | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5.1-thinking | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Chat Completions |
gpt-5.1-2025-11-13 | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Chat Completions |
gpt-5 | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5-chat | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Chat Completions |
gpt-5-mini | Usage | $0.25 / 1M tokens | $2 / 1M tokens | $0.025 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5-nano | Usage | $0.05 / 1M tokens | $0.4 / 1M tokens | $0.005 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5-pro | Usage | $15 / 1M tokens | $120 / 1M tokens | $1.5 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-5-pro-2025-10-06 | Usage | $15 / 1M tokens | $120 / 1M tokens | $1.5 / 1M tokens | — | default | Chat Completions |
gpt-5-2025-08-07 | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Chat Completions |
gpt-5-mini-2025-08-07 | Usage | $0.25 / 1M tokens | $2 / 1M tokens | $0.025 / 1M tokens | — | default | Chat Completions |
gpt-5-nano-2025-08-07 | Usage | $0.05 / 1M tokens | $0.4 / 1M tokens | $0.005 / 1M tokens | — | default | Chat Completions |
gpt-4.1 | Usage | $2 / 1M tokens | $8 / 1M tokens | $0.5 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-4.1-mini | Usage | $0.4 / 1M tokens | $1.6 / 1M tokens | $0.1 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-4.1-nano | Usage | $0.1 / 1M tokens | $0.4 / 1M tokens | $0.025 / 1M tokens | — | default | Chat Completions |
gpt-4.1-2025-04-14 | Usage | $2 / 1M tokens | $8 / 1M tokens | $0.5 / 1M tokens | — | default | Chat Completions |
gpt-4.1-mini-2025-04-14 | Usage | $0.4 / 1M tokens | $1.6 / 1M tokens | $0.1 / 1M tokens | — | default | Chat Completions |
gpt-4.1-nano-2025-04-14 | Usage | $0.1 / 1M tokens | $0.4 / 1M tokens | $0.025 / 1M tokens | — | default | Chat Completions |
gpt-4o | Usage | $2.5 / 1M tokens | $10 / 1M tokens | $1.25 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-4o-audio-preview | Usage | $2.5 / 1M tokens | $10 / 1M tokens | $1.25 / 1M tokens | — | default | Chat Completions |
gpt-4o-mini | Usage | $0.15 / 1M tokens | $0.6 / 1M tokens | $0.075 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-4o-mini-audio-preview | Usage | $2 / 1M tokens | $8 / 1M tokens | $1 / 1M tokens | — | default | Chat Completions |
gpt-4o-mini-transcribe | Usage | $1.5 / 1M tokens | $6 / 1M tokens | $0.75 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
gpt-4o-mini-tts | Usage | $1.2 / 1M tokens | $18 / 1M tokens | $0.6 / 1M tokens | — | default | Chat Completions |
gpt-4o-transcribe | Usage | $8 / 1M tokens | $16 / 1M tokens | $4 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Chat Completions |
o4-mini | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.275 / 1M tokens | — | default | Chat Completions |
text-embedding-v4 | Usage | $0.07 / 1M tokens | $0.07 / 1M tokens | — | — | default | Embeddings, Chat Completions |
o4-mini-2025-04-16 | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.275 / 1M tokens | — | default | Chat Completions |
gpt-4o-2024-11-20 | Usage | $2.5 / 1M tokens | $10 / 1M tokens | $1.25 / 1M tokens | — | default | Chat Completions |
gpt-4o-2024-08-06 | Usage | $2.5 / 1M tokens | $10 / 1M tokens | $1.25 / 1M tokens | — | default | Chat Completions |
gpt-4o-mini-2024-07-18 | Usage | $0.15 / 1M tokens | $0.6 / 1M tokens | $0.075 / 1M tokens | — | default | Chat Completions |
gpt-4o-2024-05-13 | Usage | $5 / 1M tokens | $15 / 1M tokens | $2.5 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo | Usage | $0.5 / 1M tokens | $1.5 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-0125 | Usage | $0.5 / 1M tokens | $1.5 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-0613 | Usage | $1.5 / 1M tokens | $1.95 / 1M tokens | $0.15 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-1106 | Usage | $1 / 1M tokens | $2 / 1M tokens | $0.1 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-16k | Usage | $3 / 1M tokens | $3.9 / 1M tokens | $0.3 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-16k-0613 | Usage | $3 / 1M tokens | $3.9 / 1M tokens | $0.3 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-instruct | Usage | $1.5 / 1M tokens | $1.95 / 1M tokens | $0.15 / 1M tokens | — | default | Chat Completions |
dall-e-3 | Usage | $40 / 1M tokens | $40 / 1M tokens | — | — | default | Images, Chat Completions |
o3 | Usage | $3 / 1M tokens | $12 / 1M tokens | $0.75 / 1M tokens | — | default | Chat Completions |
o3-mini | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-mini-low | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-mini-medium | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-pro | Usage | $20 / 1M tokens | $80 / 1M tokens | $5 / 1M tokens | — | default | Responses |
text-embedding-3-large | Usage | $0.13 / 1M tokens | $0.13 / 1M tokens | — | — | default | Embeddings, Chat Completions |
text-embedding-3-small | Usage | $0.02 / 1M tokens | $0.02 / 1M tokens | — | — | default | Embeddings, Chat Completions |
o3-pro-2025-06-10 | Usage | $20 / 1M tokens | $80 / 1M tokens | $5 / 1M tokens | — | default | Responses |
o3-2025-04-16 | Usage | $3 / 1M tokens | $12 / 1M tokens | $0.75 / 1M tokens | — | default | Chat Completions |
o3-mini-2025-01-31 | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-mini-2025-01-31-high | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-mini-2025-01-31-low | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-mini-2025-01-31-medium | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
gpt-image-2 | Usage | $5 / 1M tokens | $30 / 1M tokens | $1.25 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Images, Chat Completions |
gpt-image-2-all | Per call | — | — | — | $0.03 / call | default | Images, Chat Completions |
gpt-image-2-vip | Per call | — | — | — | $0.03 / call | default, GPT Image 2 Reverse | Images, Chat Completions |
sora-2 | Per call | — | — | — | $0.15 / call | Sora2Official | Chat Completions |
sora-2-character | Per call | — | — | — | $0.01 / call | default | Chat Completions |
sora-2-pro | Per call | — | — | — | $0.8 / call | Sora2Official | Chat Completions |
text-embedding-ada-002 | Usage | $0.1 / 1M tokens | $0.1 / 1M tokens | — | — | default | Embeddings, Chat Completions |
gpt-image-1.5 | Usage | $5 / 1M tokens | $32 / 1M tokens | $1.25 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Images, Chat Completions |
gpt-image-1.5-2025-12-16 | Usage | $5 / 1M tokens | $32 / 1M tokens | $1.25 / 1M tokens | — | default | Images, Chat Completions |
gpt-image-1 | Usage | $5 / 1M tokens | $40 / 1M tokens | $1.25 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Images, Chat Completions |
gpt-image-1-mini | Usage | $2 / 1M tokens | $8 / 1M tokens | $0.2 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Images, Chat Completions |
o1 | Usage | $15 / 1M tokens | $60 / 1M tokens | $7.5 / 1M tokens | — | default | Chat Completions |
o1-mini | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o1-preview | Usage | $15 / 1M tokens | $60 / 1M tokens | $7.5 / 1M tokens | — | default | Chat Completions |
o1-pro | Usage | $180 / 1M tokens | $720 / 1M tokens | $90 / 1M tokens | — | default | Chat Completions |
tts-1 | Usage | $30 / 1M tokens | $30 / 1M tokens | — | — | default | Chat Completions |
tts-1-hd | Usage | $60 / 1M tokens | $60 / 1M tokens | — | — | default | Chat Completions |
whisper-1 | Usage | $60 / 1M tokens | $0 / 1M tokens | — | — | default | Chat Completions |
o1-pro-2025-03-19 | Usage | $180 / 1M tokens | $720 / 1M tokens | $90 / 1M tokens | — | default | Chat Completions |
o1-2024-12-17 | Usage | $15 / 1M tokens | $60 / 1M tokens | $7.5 / 1M tokens | — | default | Chat Completions |
o1-mini-2024-09-12 | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o1-preview-2024-09-12 | Usage | $15 / 1M tokens | $60 / 1M tokens | $7.5 / 1M tokens | — | default | Chat Completions |
omni-moderation-latest | Usage | $0.2 / 1M tokens | $0.2 / 1M tokens | — | — | default | Chat Completions |
gpt-oss-120b | Usage | $0.5 / 1M tokens | $2 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
gpt-oss-20b | Usage | $0.1 / 1M tokens | $0.4 / 1M tokens | $0.01 / 1M tokens | — | default | Chat Completions |
sora-character | Per call | — | — | — | $0.01 / call | default | Chat Completions |
omni-moderation-2024-09-26 | Usage | $0.2 / 1M tokens | $0.2 / 1M tokens | — | — | default | Chat Completions |
Anthropic 20
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
claude-opus-5 | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-fable-5 | Usage | $10 / 1M tokens | $50 / 1M tokens | $1 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-sonnet-5 | Usage | $2 / 1M tokens | $10 / 1M tokens | $0.2 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-sonnet-5-thinking | Usage | $2 / 1M tokens | $10 / 1M tokens | $0.2 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-8 | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-8-thinking | Usage | $5 / 1M tokens | $15 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-7 | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-7-thinking | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-6 | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-6-thinking | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-sonnet-4-6 | Usage | $3 / 1M tokens | $15 / 1M tokens | $0.3 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-sonnet-4-6-thinking | Usage | $3 / 1M tokens | $15 / 1M tokens | $0.3 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-5-20251101 | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-5-20251101-thinking | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-haiku-4-5-20251001 | Usage | $1 / 1M tokens | $5 / 1M tokens | $0.1 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-haiku-4-5-20251001-thinking | Usage | $1 / 1M tokens | $5 / 1M tokens | $0.1 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-sonnet-4-5-20250929 | Usage | $3 / 1M tokens | $15 / 1M tokens | $0.3 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-sonnet-4-5-20250929-thinking | Usage | $3 / 1M tokens | $15 / 1M tokens | $0.3 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-sonnet-4-20250514 | Usage | $3 / 1M tokens | $15 / 1M tokens | $0.3 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-sonnet-4-20250514-thinking | Usage | $3 / 1M tokens | $15 / 1M tokens | $0.3 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
Google 20
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
gemini-3.6-flash | Usage | $1.5 / 1M tokens | $7.5 / 1M tokens | $0.15 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3.5-flash | Usage | $1.5 / 1M tokens | $9 / 1M tokens | $0.15 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3.5-flash-lite | Usage | $0.3 / 1M tokens | $2.502 / 1M tokens | $0.03 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3.1-flash-image | Per call | — | — | — | $0.055 / call | default | Gemini, Chat Completions |
gemini-3.1-flash-image-preview | Per call | — | — | — | $0.055 / call | default | Gemini, Chat Completions |
gemini-3.1-flash-lite | Usage | $0.25 / 1M tokens | $1.5 / 1M tokens | $0.025 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3.1-flash-lite-image | Per call | — | — | — | $0.025 / call | default | Gemini, Chat Completions |
gemini-3.1-flash-lite-preview | Usage | $0.25 / 1M tokens | $1.5 / 1M tokens | $0.025 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3.1-pro-preview † | Usage | $2 / 1M tokens | $12 / 1M tokens | $0.2 / 1M tokens | — | default | Gemini, Chat Completions |
veo-3.1-fast-generate-preview | Per call | — | — | — | $0.3 / call | default | Chat Completions |
veo-3.1-generate-preview | Per call | — | — | — | $1.2 / call | default | Chat Completions |
gemini-3-flash-preview | Usage | $0.44 / 1M tokens | $2.64 / 1M tokens | $0.044 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3-flash-preview-thinking | Usage | $0.44 / 1M tokens | $2.64 / 1M tokens | $0.044 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3-pro-image | Per call | — | — | — | $0.09 / call | default | Gemini, Chat Completions |
gemini-3-pro-image-preview | Per call | — | — | — | $0.09 / call | default | Gemini, Chat Completions |
gemini-2.5-flash | Usage | $0.3 / 1M tokens | $2.4 / 1M tokens | $0.03 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-2.5-flash-image | Per call | — | — | — | $0.02 / call | default | Gemini, Chat Completions |
gemini-2.5-flash-lite | Usage | $0.1 / 1M tokens | $0.4 / 1M tokens | $0.01 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-2.5-flash-nothinking | Usage | $0.3 / 1M tokens | $2.4 / 1M tokens | $0.03 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-2.5-pro † | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Gemini, Chat Completions |
xAI 30
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
grok-4.20-0309-non-reasoning | Usage | $1.25 / 1M tokens | $2.5 / 1M tokens | $0.2 / 1M tokens | — | default | Chat Completions |
grok-4.20-0309-reasoning | Usage | $1.25 / 1M tokens | $2.5 / 1M tokens | $0.2 / 1M tokens | — | default | Chat Completions |
grok-4.5 † | Usage | $2 / 1M tokens | $6 / 1M tokens | $0.5 / 1M tokens | — | default | Chat Completions |
grok-4.3 † | Usage | $1.25 / 1M tokens | $2.5 / 1M tokens | $0.2 / 1M tokens | — | default | Chat Completions |
grok-4-1-fast-non-reasoning-latest | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-1-fast-reasoning-latest | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-1-fast | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-1-fast-non-reasoning | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-1-fast-reasoning | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-fast-non-reasoning-latest | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-fast-reasoning-latest | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-latest | Usage | $3 / 1M tokens | $15 / 1M tokens | — | — | default | Chat Completions |
grok-4 | Usage | $3 / 1M tokens | $15 / 1M tokens | — | — | default | Chat Completions |
grok-4-0709 | Usage | $3 / 1M tokens | $15 / 1M tokens | — | — | default | Chat Completions |
grok-4-fast | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-fast-non-reasoning | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-fast-reasoning | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-3-latest | Usage | $3 / 1M tokens | $15 / 1M tokens | — | — | default | Chat Completions |
grok-3-mini-fast-latest | Usage | $0.6 / 1M tokens | $3.6 / 1M tokens | — | — | default | Chat Completions |
grok-3-mini-latest | Usage | $0.3 / 1M tokens | $1.8 / 1M tokens | — | — | default | Chat Completions |
grok-3 | Usage | $3 / 1M tokens | $15 / 1M tokens | — | — | default | Chat Completions |
grok-3-fast | Usage | $5 / 1M tokens | $25 / 1M tokens | — | — | default | Chat Completions |
grok-3-mini | Usage | $0.3 / 1M tokens | $1.8 / 1M tokens | — | — | default | Chat Completions |
grok-3-mini-beta | Usage | $0.3 / 1M tokens | $1.8 / 1M tokens | — | — | default | Chat Completions |
grok-3-mini-fast | Usage | $0.6 / 1M tokens | $3.6 / 1M tokens | — | — | default | Chat Completions |
grok-3-mini-fast-beta | Usage | $0.6 / 1M tokens | $3.6 / 1M tokens | — | — | default | Chat Completions |
grok-2-vision-latest | Usage | $2 / 1M tokens | $10 / 1M tokens | — | — | default | Chat Completions |
grok-2-vision | Usage | $2 / 1M tokens | $10 / 1M tokens | — | — | default | Chat Completions |
grok-2-vision-1212 | Usage | $2 / 1M tokens | $10 / 1M tokens | — | — | default | Chat Completions |
grok-build-0.1 | Usage | $1 / 1M tokens | $2 / 1M tokens | $0.2 / 1M tokens | — | default | Chat Completions |
DeepSeek 7
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
deepseek-v4-flash | Usage | $0.14 / 1M tokens | $0.28 / 1M tokens | $0.0028 / 1M tokens | — | default | Anthropic Messages, Chat Completions |
deepseek-v4-pro | Usage | $1.74 / 1M tokens | $3.48 / 1M tokens | $0.0145 / 1M tokens | — | default | Anthropic Messages, Chat Completions |
DeepSeek-V3.2-Exp-nothinking | Usage | $0.3 / 1M tokens | $0.45 / 1M tokens | — | — | default | Chat Completions |
DeepSeek-V3.2-Exp-thinking | Usage | $0.3 / 1M tokens | $0.45 / 1M tokens | — | — | default | Chat Completions |
deepseek-v3.2 | Usage | $0.28 / 1M tokens | $0.42 / 1M tokens | — | — | default | Chat Completions |
deepseek-v3.2-thinking | Usage | $0.28 / 1M tokens | $0.42 / 1M tokens | — | — | default | Chat Completions |
deepseek-r1 | Usage | $0.57 / 1M tokens | $2.28 / 1M tokens | $0.1425 / 1M tokens | — | default | Chat Completions |
阿里巴巴 77
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
qwen3.6-flash † | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | default | Chat Completions |
qwen3.6-max-preview † | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | default | Chat Completions |
qwen3.6-plus † | Usage | $0.3 / 1M tokens | $1.8 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-122b-a10b | Usage | $0.12 / 1M tokens | $0.96 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-27b | Usage | $0.09 / 1M tokens | $0.72 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-35b-a3b | Usage | $0.06 / 1M tokens | $0.48 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-397b-a17b | Usage | $0.18 / 1M tokens | $1.08 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-flash † | Usage | $0.03 / 1M tokens | $0.3 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-plus † | Usage | $0.114 / 1M tokens | $0.684 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-flash-2026-02-23 † | Usage | $0.03 / 1M tokens | $0.3 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-plus-2026-02-15 † | Usage | $0.114 / 1M tokens | $0.684 / 1M tokens | — | — | default | Chat Completions |
qwen3-235b-a22b | Usage | $1 / 1M tokens | $10 / 1M tokens | — | — | default | Chat Completions |
qwen3-30b-a3b | Usage | $0.2 / 1M tokens | $2 / 1M tokens | — | — | default | Chat Completions |
qwen3-32b | Usage | $0.4 / 1M tokens | $4 / 1M tokens | — | — | default | Chat Completions |
qwen3-coder-480b-a35b-instruct † | Usage | $3 / 1M tokens | $15 / 1M tokens | — | — | default | Chat Completions |
qwen3-coder-flash † | Usage | $0.5 / 1M tokens | $5 / 1M tokens | — | — | default | Chat Completions |
qwen3-coder-plus † | Usage | $5 / 1M tokens | $20 / 1M tokens | — | — | default | Chat Completions |
qwen3-max † | Usage | $1.2 / 1M tokens | $6 / 1M tokens | — | — | default | Chat Completions |
qwen3-max-preview † | Usage | $1.2 / 1M tokens | $6 / 1M tokens | — | — | default | Chat Completions |
qwen3-next-80b-a3b-instruct | Usage | $0.15 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
qwen3-omni-flash | Usage | $2 / 1M tokens | $20 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-235b-a22b-instruct | Usage | $0.3 / 1M tokens | $3 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-235b-a22b-thinking | Usage | $0.3 / 1M tokens | $3 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-30b-a3b-thinking | Usage | $0.12 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-32b-thinking | Usage | $0.3 / 1M tokens | $3 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-embedding | Usage | $0.25 / 1M tokens | $0.25 / 1M tokens | — | — | default | Embeddings, Chat Completions |
qwen3-vl-flash | Usage | $0.1 / 1M tokens | $0.8 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-plus † | Usage | $0.3 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-flash-2025-10-15 | Usage | $0.1 / 1M tokens | $0.8 / 1M tokens | — | — | default | Chat Completions |
qwen3-coder-plus-2025-09-23 † | Usage | $2 / 1M tokens | $20 / 1M tokens | — | — | default | Chat Completions |
qwen3-max-2025-09-23 † | Usage | $1.2 / 1M tokens | $6 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-plus-2025-09-23 † | Usage | $0.3 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
qwen3-omni-flash-2025-09-15 | Usage | $2 / 1M tokens | $20 / 1M tokens | — | — | default | Chat Completions |
qwen3-coder-plus-2025-07-22 † | Usage | $2 / 1M tokens | $20 / 1M tokens | — | — | default | Chat Completions |
qwen3-235b-a22b-instruct-2507 | Usage | $1 / 1M tokens | $10 / 1M tokens | — | — | default | Chat Completions |
qwen3-235b-a22b-thinking-2507 | Usage | $1.6 / 1M tokens | $12.8 / 1M tokens | — | — | default | Chat Completions |
qwen3-30b-a3b-instruct-2507 | Usage | $0.2 / 1M tokens | $0.8 / 1M tokens | — | — | default | Chat Completions |
qwen3-30b-a3b-thinking-2507 | Usage | $0.2 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
wan2.7-i2v | Usage | $1.2 / 1M tokens | $1.2 / 1M tokens | — | — | Wan | Chat Completions |
wan2.7-r2v | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | Wan | Chat Completions |
wan2.7-t2v | Usage | $1.2 / 1M tokens | $1.2 / 1M tokens | — | — | Wan | Chat Completions |
wan2.7-videoedit | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | Wan | Chat Completions |
wan2.6-i2v | Usage | $1.2 / 1M tokens | $1.2 / 1M tokens | — | — | Wan | Chat Completions |
wan2.6-r2v | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | Wan | Chat Completions |
wan2.6-r2v-flash | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | Wan | Chat Completions |
wan2.6-t2v | Usage | $1.2 / 1M tokens | $1.2 / 1M tokens | — | — | Wan | Chat Completions |
qwen2-72b-instruct | Usage | $2.8572 / 1M tokens | $2.8572 / 1M tokens | — | — | default | Chat Completions |
multimodal-embedding-v1 | Usage | $0 / 1M tokens | $0 / 1M tokens | — | — | default | Embeddings, Chat Completions |
qvq-max-latest | Usage | $1.2 / 1M tokens | $4.8 / 1M tokens | — | — | default | Chat Completions |
qwen-max-latest | Usage | $1.6 / 1M tokens | $6.4 / 1M tokens | — | — | default | Chat Completions |
qwen-plus-latest † | Usage | $0.4 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
qwen-turbo-latest | Usage | $0.2 / 1M tokens | $0.6 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-max-latest | Usage | $0.2 / 1M tokens | $0.8 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-ocr-latest | Usage | $0.044 / 1M tokens | $0.07348 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-plus-latest | Usage | $0.2 / 1M tokens | $0.6 / 1M tokens | — | — | default | Chat Completions |
qwq-plus-latest | Usage | $0.8 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
qvq-max | Usage | $1.2 / 1M tokens | $4.8 / 1M tokens | — | — | default | Chat Completions |
qvq-plus | Usage | $0.28 / 1M tokens | $0.7 / 1M tokens | — | — | default | Chat Completions |
qwen-long | Usage | $0.07 / 1M tokens | $0.28 / 1M tokens | — | — | default | Chat Completions |
qwen-max | Usage | $1.6 / 1M tokens | $6.4 / 1M tokens | — | — | default | Chat Completions |
qwen-max-longcontext | Usage | $3.2 / 1M tokens | $3.2 / 1M tokens | — | — | default | Chat Completions |
qwen-mt-plus | Usage | $2 / 1M tokens | $8 / 1M tokens | — | — | default | Chat Completions |
qwen-mt-turbo | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
qwen-plus † | Usage | $0.4 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
qwen-turbo | Usage | $0.2 / 1M tokens | $0.6 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-max | Usage | $0.2 / 1M tokens | $0.8 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-ocr | Usage | $0.72 / 1M tokens | $0.72 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-plus | Usage | $0.2 / 1M tokens | $0.6 / 1M tokens | — | — | default | Chat Completions |
qwq-32b | Usage | $0.4 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
qwq-plus | Usage | $0.8 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-ocr-2025-11-20 | Usage | $0.044 / 1M tokens | $0.07348 / 1M tokens | — | — | default | Chat Completions |
qwen-plus-2025-09-11 † | Usage | $0.4 / 1M tokens | $4 / 1M tokens | — | — | default | Chat Completions |
qwen-turbo-2025-07-15 | Usage | $0.2 / 1M tokens | $1.6 / 1M tokens | — | — | default | Chat Completions |
qwen-plus-2025-07-14 | Usage | $0.4 / 1M tokens | $4 / 1M tokens | — | — | default | Chat Completions |
qvq-max-2025-05-15 | Usage | $1 / 1M tokens | $4 / 1M tokens | — | — | default | Chat Completions |
qvq-plus-2025-05-15 | Usage | $0.28 / 1M tokens | $0.7 / 1M tokens | — | — | default | Chat Completions |
qwq-plus-2025-03-05 | Usage | $0.8 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
字节跳动 10
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
seedream-5-0-260128 | Per call | — | — | — | $0.035 / call | default | Chat Completions |
seedream-4-5-251128 | Per call | — | — | — | $0.045 / call | default | Chat Completions |
seedream-4-0-250828 | Per call | — | — | — | $0.035 / call | default | Chat Completions |
doubao-seedance-2-0-mini-260615 | Usage | $23 / 1M tokens | $23 / 1M tokens | — | — | SeeDance2 | Chat Completions |
seed-2-0-lite-260428 | Usage | $0.25 / 1M tokens | $2 / 1M tokens | — | — | default | Chat Completions |
seed-2-0-mini-260428 | Usage | $0.1 / 1M tokens | $0.4 / 1M tokens | — | — | default | Chat Completions |
seed-2-0-code-preview-260328 | Usage | $0.5 / 1M tokens | $3 / 1M tokens | — | — | default | Chat Completions |
seed-2-0-pro-260328 † | Usage | $0.5 / 1M tokens | $3 / 1M tokens | — | — | default | Chat Completions |
doubao-seedance-2-0-260128 | Usage | $46 / 1M tokens | $46 / 1M tokens | — | — | SeeDance2 | Chat Completions |
doubao-seedance-2-0-fast-260128 | Usage | $37 / 1M tokens | $37 / 1M tokens | — | — | SeeDance2 | Chat Completions |
智谱 10
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
glm-5.2 | Usage | $1.142 / 1M tokens | $3.997 / 1M tokens | $0.2284 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
glm-5.1 † | Usage | $0.84 / 1M tokens | $3.36 / 1M tokens | $0.084 / 1M tokens | — | default | Chat Completions |
glm-5 † | Usage | $0.56 / 1M tokens | $2.52 / 1M tokens | $0.056 / 1M tokens | — | default | Chat Completions |
glm-4.7 | Usage | $0.6 / 1M tokens | $2.16 / 1M tokens | $0.06 / 1M tokens | — | default | Chat Completions |
glm-4.6 | Usage | $0.5 / 1M tokens | $2 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
glm-4.6v | Usage | $0.28 / 1M tokens | $0.84 / 1M tokens | $0.028 / 1M tokens | — | default | Chat Completions |
glm-4.5 | Usage | $0.5 / 1M tokens | $2 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
glm-4.5-air | Usage | $0.2 / 1M tokens | $1 / 1M tokens | $0.02 / 1M tokens | — | default | Chat Completions |
glm-4.5-flash | Usage | $0.01 / 1M tokens | $0.04 / 1M tokens | $0.001 / 1M tokens | — | default | Chat Completions |
glm-4.5v | Usage | $0.5 / 1M tokens | $1.5 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
Moonshot 5
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
kimi-k2.6 | Usage | $0.95 / 1M tokens | $4 / 1M tokens | $0.16 / 1M tokens | — | default | Chat Completions |
kimi-k2.5 | Usage | $0.6 / 1M tokens | $3.15 / 1M tokens | $0.1 / 1M tokens | — | default | Chat Completions |
kimi-k2 | Usage | $0.56 / 1M tokens | $2.24 / 1M tokens | $0.056 / 1M tokens | — | default | Chat Completions |
kimi-k2-128k | Usage | $0.56 / 1M tokens | $2.24 / 1M tokens | $0.056 / 1M tokens | — | default | Chat Completions |
kimi-k2-instruct | Usage | $0.56 / 1M tokens | $0.56 / 1M tokens | $0.056 / 1M tokens | — | default | Chat Completions |
MiniMax 2
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
MiniMax-M2.5 | Usage | $0.3 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
MiniMax-M2.1 | Usage | $0.3 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
Black Forest Labs 5
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
flux-2-flex | Per call | — | — | — | $0.06 / call | default | Images, Chat Completions |
flux-2-max | Per call | — | — | — | $0.07 / call | default | Images, Chat Completions |
flux-2-pro | Per call | — | — | — | $0.03 / call | default | Images, Chat Completions |
flux-kontext-max | Per call | — | — | — | $0.07 / call | default | Images, Chat Completions |
flux-kontext-pro | Per call | — | — | — | $0.035 / call | default | Images, Chat Completions |
美团 1
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
longcat-flash-chat | Usage | $0.25 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
阶跃星辰 1
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
step-3.5-flash | Usage | $0.1 / 1M tokens | $0.3 / 1M tokens | — | — | default | Chat Completions |
Tiered pricing
Models marked with † use different price tiers based on the token count of a single request. The table below shows the input and output price for each tier. Confirm the final token count and charge in call logs.
| Model ID | Tokens per request | Input price | Output price |
|---|---|---|---|
qwen3.6-flash | 0–262,144 | $75 / 1M tokens | $450 / 1M tokens |
qwen3.6-flash | 262,145–1,024,000 | $300 / 1M tokens | $1800 / 1M tokens |
qwen3.6-max-preview | 0–131,072 | $75 / 1M tokens | $450 / 1M tokens |
qwen3.6-max-preview | 131,073–262,144 | $124.22 / 1M tokens | $745.31 / 1M tokens |
qwen3.6-plus | 0–262,144 | $0.3 / 1M tokens | $1.8 / 1M tokens |
qwen3.6-plus | 262,145–1,024,000 | $1.2 / 1M tokens | $7.2 / 1M tokens |
qwen3.5-flash | 0–131,072 | $0.03 / 1M tokens | $0.3 / 1M tokens |
qwen3.5-flash | 131,073–262,144 | $0.122143 / 1M tokens | $1.2214 / 1M tokens |
qwen3.5-flash | 262,145 and above | $0.182143 / 1M tokens | $1.8214 / 1M tokens |
qwen3.5-plus | 0–131,072 | $0.114 / 1M tokens | $0.684 / 1M tokens |
qwen3.5-plus | 131,073–262,144 | $0.28 / 1M tokens | $1.68 / 1M tokens |
qwen3.5-plus | 262,145 and above | $0.56 / 1M tokens | $3.36 / 1M tokens |
qwen3.5-flash-2026-02-23 | 0–131,072 | $0.03 / 1M tokens | $0.3 / 1M tokens |
qwen3.5-flash-2026-02-23 | 131,073–262,144 | $0.122143 / 1M tokens | $1.2214 / 1M tokens |
qwen3.5-flash-2026-02-23 | 262,145 and above | $0.182143 / 1M tokens | $1.8214 / 1M tokens |
qwen3.5-plus-2026-02-15 | 0–131,072 | $0.114 / 1M tokens | $0.684 / 1M tokens |
qwen3.5-plus-2026-02-15 | 131,073–262,144 | $0.28 / 1M tokens | $1.68 / 1M tokens |
qwen3.5-plus-2026-02-15 | 262,145 and above | $0.56 / 1M tokens | $3.36 / 1M tokens |
qwen3-coder-480b-a35b-instruct | 0–32,768 | $3 / 1M tokens | $15 / 1M tokens |
qwen3-coder-480b-a35b-instruct | 32,769–131,072 | $5.4 / 1M tokens | $27 / 1M tokens |
qwen3-coder-480b-a35b-instruct | 131,073–204,800 | $9 / 1M tokens | $45 / 1M tokens |
qwen3-coder-flash — includes a model-specific price, about 50% of the official nominal price | 0–32,000 | $0.5 / 1M tokens | $2 / 1M tokens |
qwen3-coder-flash | 32,001–128,000 | $0.75 / 1M tokens | $3 / 1M tokens |
qwen3-coder-flash | 128,001–256,000 | $1.25 / 1M tokens | $5 / 1M tokens |
qwen3-coder-flash | 256,001–1,000,000 | $2.5 / 1M tokens | $12.5 / 1M tokens |
qwen3-coder-plus | 0–32,000 | $5 / 1M tokens | $20 / 1M tokens |
qwen3-coder-plus | 32,001–128,000 | $7.5 / 1M tokens | $30 / 1M tokens |
qwen3-coder-plus | 128,001–256,000 | $12.5 / 1M tokens | $50 / 1M tokens |
qwen3-coder-plus | 256,001–1,000,000 | $25 / 1M tokens | $250 / 1M tokens |
qwen3-max — includes a model-specific price, about 48% of the official nominal price | 0–32,000 | $1.2 / 1M tokens | $4.8 / 1M tokens |
qwen3-max | 32,001–128,000 | $1.92 / 1M tokens | $7.68 / 1M tokens |
qwen3-max | 128,001–256,000 | $3.36 / 1M tokens | $13.44 / 1M tokens |
qwen3-max-preview — includes a model-specific price, about 20% of the official nominal price | 0–32,000 | $1.2 / 1M tokens | $4.8 / 1M tokens |
qwen3-max-preview | 32,001–128,000 | $2 / 1M tokens | $8 / 1M tokens |
qwen3-max-preview | 128,001–256,000 | $3 / 1M tokens | $12 / 1M tokens |
qwen3-vl-plus — includes a model-specific price, about 30% of the official nominal price | 0–32,000 | $0.3 / 1M tokens | $3 / 1M tokens |
qwen3-vl-plus | 32,001–128,000 | $0.45 / 1M tokens | $4.5 / 1M tokens |
qwen3-vl-plus | 128,001–256,000 | $0.9 / 1M tokens | $9 / 1M tokens |
qwen3-coder-plus-2025-09-23 — includes a model-specific price, about 50% of the official nominal price | 0–32,000 | $2 / 1M tokens | $8 / 1M tokens |
qwen3-coder-plus-2025-09-23 | 32,001–128,000 | $3 / 1M tokens | $12 / 1M tokens |
qwen3-coder-plus-2025-09-23 | 128,001–256,000 | $5 / 1M tokens | $20 / 1M tokens |
qwen3-coder-plus-2025-09-23 | 256,001–1,000,000 | $10 / 1M tokens | $100 / 1M tokens |
qwen3-max-2025-09-23 — includes a model-specific price, about 20% of the official nominal price | 0–32,000 | $1.2 / 1M tokens | $4.8 / 1M tokens |
qwen3-max-2025-09-23 | 32,001–128,000 | $2 / 1M tokens | $8 / 1M tokens |
qwen3-max-2025-09-23 | 128,001–256,000 | $3 / 1M tokens | $12 / 1M tokens |
qwen3-vl-plus-2025-09-23 — includes a model-specific price, about 30% of the official nominal price | 0–32,000 | $0.3 / 1M tokens | $3 / 1M tokens |
qwen3-vl-plus-2025-09-23 | 32,001–128,000 | $0.45 / 1M tokens | $4.5 / 1M tokens |
qwen3-vl-plus-2025-09-23 | 128,001–256,000 | $0.9 / 1M tokens | $9 / 1M tokens |
qwen3-coder-plus-2025-07-22 — includes a model-specific price, about 50% of the official nominal price | 0–32,000 | $2 / 1M tokens | $8 / 1M tokens |
qwen3-coder-plus-2025-07-22 | 32,001–128,000 | $3 / 1M tokens | $12 / 1M tokens |
qwen3-coder-plus-2025-07-22 | 128,001–256,000 | $5 / 1M tokens | $20 / 1M tokens |
qwen3-coder-plus-2025-07-22 | 256,001–1,000,000 | $10 / 1M tokens | $100 / 1M tokens |
qwen-plus-latest — includes a model-specific price, about 50% of the official nominal price | 0–128,000 | $0.4 / 1M tokens | $1 / 1M tokens |
qwen-plus-latest | 128,001–256,000 | $1.2 / 1M tokens | $10 / 1M tokens |
qwen-plus-latest | 256,001–1,000,000 | $2.4 / 1M tokens | $24 / 1M tokens |
qwen-plus — includes a model-specific price, about 50% of the official nominal price | 0–128,000 | $0.4 / 1M tokens | $1 / 1M tokens |
qwen-plus | 128,001–256,000 | $1.2 / 1M tokens | $10 / 1M tokens |
qwen-plus | 256,001–1,000,000 | $2.4 / 1M tokens | $24 / 1M tokens |
qwen-plus-2025-09-11 — includes a model-specific price, about 50% of the official nominal price | 0–128,000 | $0.4 / 1M tokens | $1 / 1M tokens |
qwen-plus-2025-09-11 | 128,001–256,000 | $1.2 / 1M tokens | $10 / 1M tokens |
qwen-plus-2025-09-11 | 256,001–1,000,000 | $2.4 / 1M tokens | $24 / 1M tokens |
glm-5.1 | 0–32,768 | $0.84 / 1M tokens | $3.36 / 1M tokens |
glm-5.1 | 32,769 and above | $1.14 / 1M tokens | $3.99 / 1M tokens |
glm-5 | 0–32,000 | $0.56 / 1M tokens | $2.52 / 1M tokens |
glm-5 | 32,001 and above | $0.86 / 1M tokens | $3.096 / 1M tokens |
seed-2-0-pro-260328 | 0–131,072 | $0.5 / 1M tokens | $3 / 1M tokens |
seed-2-0-pro-260328 | 131,073–262,144 | $1 / 1M tokens | $6 / 1M tokens |
gemini-3.1-pro-preview | 0–200,000 | $2 / 1M tokens | $12 / 1M tokens |
gemini-3.1-pro-preview | 200,001 and above | $4 / 1M tokens | $18 / 1M tokens |
gemini-2.5-pro | 0–200,000 | $1.25 / 1M tokens | $10 / 1M tokens |
gemini-2.5-pro | 200,001 and above | $2.5 / 1M tokens | $15 / 1M tokens |
gpt-5.6-luna | 0–272,000 | $1 / 1M tokens | $6 / 1M tokens |
gpt-5.6-luna | 272,001 and above | $2 / 1M tokens | $9 / 1M tokens |
gpt-5.6-sol | 0–272,000 | $5 / 1M tokens | $30 / 1M tokens |
gpt-5.6-sol | 272,001 and above | $10 / 1M tokens | $45 / 1M tokens |
gpt-5.6-terra | 0–272,000 | $2.5 / 1M tokens | $15 / 1M tokens |
gpt-5.6-terra | 272,001 and above | $5 / 1M tokens | $22.5 / 1M tokens |
gpt-5.5 | 0–278,528 | $5 / 1M tokens | $30 / 1M tokens |
gpt-5.5 | 278,529 and above | $10 / 1M tokens | $45 / 1M tokens |
gpt-5.5-pro | 0–278,528 | $30 / 1M tokens | $180 / 1M tokens |
gpt-5.5-pro | 278,529 and above | $60 / 1M tokens | $270 / 1M tokens |
gpt-5.4 | 0–278,528 | $2.5 / 1M tokens | $15 / 1M tokens |
gpt-5.4 | 278,529 and above | $5 / 1M tokens | $22.5 / 1M tokens |
gpt-5.4-pro | 0–278,528 | $30 / 1M tokens | $180 / 1M tokens |
gpt-5.4-pro | 278,529 and above | $60 / 1M tokens | $270 / 1M tokens |
grok-4.5 | 0–204,800 | $2 / 1M tokens | $6 / 1M tokens |
grok-4.5 | 204,801 and above | $4 / 1M tokens | $12 / 1M tokens |
grok-4.3 | 0–204,800 | $1.25 / 1M tokens | $2.5 / 1M tokens |
grok-4.3 | 204,801 and above | $2.5 / 1M tokens | $5 / 1M tokens |
Frequently asked questions
Is this price read in real time?
This page is generated periodically from current pricing configuration and shows its update time. Model changes, account groups, or dedicated contracts can produce a different console price; confirm final charges in the console and call logs.
When does the cache-read price apply?
Only when the request hits a supported prompt cache. If the model has no cache setting, the cache is missed, or the calling method does not support caching, normal input pricing applies.
How does tiered pricing work?
The system selects a tier from the token count of each request. The tier table lists both input and output prices; confirm final usage and charges in call logs.
Does a US$700 top-up change the listed model price?
No. A qualifying single top-up receives a 5% balance bonus while the default model list price remains unchanged. Confirm the credited balance and charges in the console.
How do I call a model after finding its ID?
Confirm the endpoint and token group shown in the table, then create the matching token in the console. OpenAI-compatible models generally use /v1/chat/completions or /v1/responses; Gemini, image, and specialized models should follow their dedicated documentation.
Pricing snapshot generated: 2026-07-25T14:52:31.341Z