Model Catalog and Pricing
This page lists 249 online LaoZhang API models with input, output, cache-read, per-call, and tiered prices, token groups, and supported endpoints. The data was updated on 2026-09-02, 12:12 (UTC+8). Confirm final availability and charges in the signed-in console and call logs.
Online models, cache-read prices, and tiered pricing rules are refreshed from current system configuration. Use the signed-in console to confirm account groups and current prices, and call logs to confirm actual charges.
Pricing notes
All amounts on this page are shown in US dollars. Usage-based input, output, and cache-read prices are per million tokens ($/1M). Per-call models are priced for each call. The table shows default list prices; account contracts or dedicated routes may use different prices.
Enterprise purchasing and discounts
Contact the site owner or support for enterprise purchasing, volume usage, contracts, invoices, discounts, or special payment arrangements. The documentation does not publish fixed discount thresholds or percentages. Final commercial terms, credited balance, and charges follow the confirmed arrangement, console, and call logs.
A model can support multiple token groups with different routes, permissions, or billing modes. Confirm the groups and actual prices available to your account in the console.
Field reference
| Field | Meaning |
|---|---|
| Model ID | The model name used in API requests. A † mark identifies a model with token-volume pricing tiers. |
| Input / output | The usage-based US dollar price per 1M tokens. |
| Cache read | The input price when a supported prompt cache is hit. “—” means no current cache-read price is listed. |
| Per-call price | The US dollar price per call. A call can mean a request, image, or task depending on the model documentation. |
| Token groups | Available group names. Account access and actual prices require a console check. |
| Supported endpoints | Declared protocol entry points. This does not guarantee that every client parameter is fully compatible. |
Online model catalog
Models within each vendor are ordered from newer to older versions. Use your browser find command or the search box above. Wide tables scroll horizontally on small screens.
OpenAI 92
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
gpt-5.6-luna † | Usage | $0.2 / 1M tokens | $1.2 / 1M tokens | $0.02 / 1M tokens | — | default | Chat Completions |
gpt-5.6-sol † | Usage | $5 / 1M tokens | $30 / 1M tokens | $0.5 / 1M tokens | — | default | Chat Completions |
gpt-5.6-terra † | Usage | $2 / 1M tokens | $12 / 1M tokens | $0.2 / 1M tokens | — | default | Chat Completions |
gpt-5.5 † | Usage | $5 / 1M tokens | $30 / 1M tokens | $0.5 / 1M tokens | — | default | Chat Completions |
gpt-5.4 † | Usage | $2.5 / 1M tokens | $15 / 1M tokens | $0.25 / 1M tokens | — | default | Chat Completions |
gpt-5.4-mini | Usage | $0.75 / 1M tokens | $4.5 / 1M tokens | $0.075 / 1M tokens | — | default | Chat Completions |
gpt-5.4-nano | Usage | $0.2 / 1M tokens | $1.25 / 1M tokens | $0.02 / 1M tokens | — | default | Chat Completions |
gpt-5.4-pro † | Usage | $30 / 1M tokens | $180 / 1M tokens | $3 / 1M tokens | — | default | Chat Completions |
gpt-5.3 | Usage | $75 / 1M tokens | $450 / 1M tokens | $7.5 / 1M tokens | — | default | Chat Completions |
gpt-5.2 | Usage | $1.75 / 1M tokens | $14 / 1M tokens | $0.175 / 1M tokens | — | default | Chat Completions |
gpt-5.2-2025-12-11 | Usage | $1.75 / 1M tokens | $14 / 1M tokens | $0.175 / 1M tokens | — | default | Chat Completions |
gpt-5.1 | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Chat Completions |
gpt-5.1-thinking | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Chat Completions |
gpt-5.1-2025-11-13 | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Chat Completions |
gpt-5 | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Chat Completions |
gpt-5-chat | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Chat Completions |
gpt-5-mini | Usage | $0.25 / 1M tokens | $2 / 1M tokens | $0.025 / 1M tokens | — | default | Chat Completions |
gpt-5-nano | Usage | $0.05 / 1M tokens | $0.4 / 1M tokens | $0.005 / 1M tokens | — | default | Chat Completions |
gpt-5-pro | Usage | $15 / 1M tokens | $120 / 1M tokens | $1.5 / 1M tokens | — | default | Chat Completions |
gpt-5-pro-2025-10-06 | Usage | $15 / 1M tokens | $120 / 1M tokens | $1.5 / 1M tokens | — | default | Chat Completions |
gpt-5-2025-08-07 | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Chat Completions |
gpt-5-mini-2025-08-07 | Usage | $0.25 / 1M tokens | $2 / 1M tokens | $0.025 / 1M tokens | — | default | Chat Completions |
gpt-5-nano-2025-08-07 | Usage | $0.05 / 1M tokens | $0.4 / 1M tokens | $0.005 / 1M tokens | — | default | Chat Completions |
gpt-4.1 | Usage | $2 / 1M tokens | $8 / 1M tokens | $0.5 / 1M tokens | — | default | Chat Completions |
gpt-4.1-mini | Usage | $0.4 / 1M tokens | $1.6 / 1M tokens | $0.1 / 1M tokens | — | default | Chat Completions |
gpt-4.1-nano | Usage | $0.1 / 1M tokens | $0.4 / 1M tokens | $0.025 / 1M tokens | — | default | Chat Completions |
gpt-4.1-2025-04-14 | Usage | $2 / 1M tokens | $8 / 1M tokens | $0.5 / 1M tokens | — | default | Chat Completions |
gpt-4.1-mini-2025-04-14 | Usage | $0.4 / 1M tokens | $1.6 / 1M tokens | $0.1 / 1M tokens | — | default | Chat Completions |
gpt-4.1-nano-2025-04-14 | Usage | $0.1 / 1M tokens | $0.4 / 1M tokens | $0.025 / 1M tokens | — | default | Chat Completions |
gpt-4o | Usage | $2.5 / 1M tokens | $10 / 1M tokens | $1.25 / 1M tokens | — | default | Chat Completions |
gpt-4o-audio-preview | Usage | $2.5 / 1M tokens | $10 / 1M tokens | $1.25 / 1M tokens | — | default | Chat Completions |
gpt-4o-mini | Usage | $0.15 / 1M tokens | $0.6 / 1M tokens | $0.075 / 1M tokens | — | default | Chat Completions |
gpt-4o-mini-audio-preview | Usage | $2 / 1M tokens | $8 / 1M tokens | $1 / 1M tokens | — | default | Chat Completions |
gpt-4o-mini-transcribe | Usage | $1.5 / 1M tokens | $6 / 1M tokens | $0.75 / 1M tokens | — | default | Chat Completions |
gpt-4o-mini-tts | Usage | $1.2 / 1M tokens | $18 / 1M tokens | $0.6 / 1M tokens | — | default | Chat Completions |
gpt-4o-transcribe | Usage | $8 / 1M tokens | $16 / 1M tokens | $4 / 1M tokens | — | default | Chat Completions |
o4-mini | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.275 / 1M tokens | — | default | Chat Completions |
text-embedding-v4 | Usage | $0.07 / 1M tokens | $0.07 / 1M tokens | — | — | default | Embeddings, Chat Completions |
o4-mini-2025-04-16 | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.275 / 1M tokens | — | default | Chat Completions |
gpt-4o-2024-11-20 | Usage | $2.5 / 1M tokens | $10 / 1M tokens | $1.25 / 1M tokens | — | default | Chat Completions |
gpt-4o-2024-08-06 | Usage | $2.5 / 1M tokens | $10 / 1M tokens | $1.25 / 1M tokens | — | default | Chat Completions |
gpt-4o-mini-2024-07-18 | Usage | $0.15 / 1M tokens | $0.6 / 1M tokens | $0.075 / 1M tokens | — | default | Chat Completions |
gpt-4o-2024-05-13 | Usage | $5 / 1M tokens | $15 / 1M tokens | $2.5 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo | Usage | $0.5 / 1M tokens | $1.5 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-0125 | Usage | $0.5 / 1M tokens | $1.5 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-0613 | Usage | $1.5 / 1M tokens | $1.95 / 1M tokens | $0.15 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-1106 | Usage | $1 / 1M tokens | $2 / 1M tokens | $0.1 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-16k | Usage | $3 / 1M tokens | $3.9 / 1M tokens | $0.3 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-16k-0613 | Usage | $3 / 1M tokens | $3.9 / 1M tokens | $0.3 / 1M tokens | — | default | Chat Completions |
gpt-3.5-turbo-instruct | Usage | $1.5 / 1M tokens | $1.95 / 1M tokens | $0.15 / 1M tokens | — | default | Chat Completions |
dall-e-3 | Usage | $40 / 1M tokens | $40 / 1M tokens | — | — | default | Images, Chat Completions |
o3 | Usage | $3 / 1M tokens | $12 / 1M tokens | $0.75 / 1M tokens | — | default | Chat Completions |
o3-mini | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-mini-low | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-mini-medium | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-pro | Usage | $20 / 1M tokens | $80 / 1M tokens | $5 / 1M tokens | — | default | Responses |
text-embedding-3-large | Usage | $0.13 / 1M tokens | $0.13 / 1M tokens | — | — | default | Embeddings, Chat Completions |
text-embedding-3-small | Usage | $0.02 / 1M tokens | $0.02 / 1M tokens | — | — | default | Embeddings, Chat Completions |
o3-pro-2025-06-10 | Usage | $20 / 1M tokens | $80 / 1M tokens | $5 / 1M tokens | — | default | Responses |
o3-2025-04-16 | Usage | $3 / 1M tokens | $12 / 1M tokens | $0.75 / 1M tokens | — | default | Chat Completions |
o3-mini-2025-01-31 | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-mini-2025-01-31-high | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-mini-2025-01-31-low | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o3-mini-2025-01-31-medium | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
gpt-image-2 | Usage | $5 / 1M tokens | $30 / 1M tokens | $1.25 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, GPT Image 2 Reverse, Sora2Official | Images, Chat Completions |
gpt-image-2-all | Per call | — | — | — | $0.03 / call | default, GPT Image 2 Reverse | Images, Chat Completions |
gpt-image-2-vip | Per call | — | — | — | $0.03 / call | default, GPT Image 2 Reverse | Images, Chat Completions |
gpt-image-2-web | Per call | — | — | — | $0.03 / call | default, GPT Image 2 Reverse | Images, Chat Completions |
sora-2 | Per call | — | — | — | $0.15 / call | Sora2Official | Chat Completions |
sora-2-character | Per call | — | — | — | $0.01 / call | default | Chat Completions |
sora-2-pro | Per call | — | — | — | $0.8 / call | Sora2Official | Chat Completions |
text-embedding-ada-002 | Usage | $0.1 / 1M tokens | $0.1 / 1M tokens | — | — | default | Embeddings, Chat Completions |
gpt-image-1.5 | Usage | $5 / 1M tokens | $32 / 1M tokens | $1.25 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Images, Chat Completions |
gpt-image-1.5-2025-12-16 | Usage | $5 / 1M tokens | $32 / 1M tokens | $1.25 / 1M tokens | — | default | Images, Chat Completions |
gpt-image-1 | Usage | $5 / 1M tokens | $40 / 1M tokens | $1.25 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Images, Chat Completions |
gpt-image-1-mini | Usage | $2 / 1M tokens | $8 / 1M tokens | $0.2 / 1M tokens | — | GPTImage2 Sora2 Enterprise, default, Sora2Official | Images, Chat Completions |
o1 | Usage | $15 / 1M tokens | $60 / 1M tokens | $7.5 / 1M tokens | — | default | Chat Completions |
o1-mini | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o1-preview | Usage | $15 / 1M tokens | $60 / 1M tokens | $7.5 / 1M tokens | — | default | Chat Completions |
o1-pro | Usage | $180 / 1M tokens | $720 / 1M tokens | $90 / 1M tokens | — | default | Chat Completions |
tts-1 | Usage | $30 / 1M tokens | $30 / 1M tokens | — | — | default | Chat Completions |
tts-1-hd | Usage | $60 / 1M tokens | $60 / 1M tokens | — | — | default | Chat Completions |
whisper-1 | Usage | $60 / 1M tokens | $0 / 1M tokens | — | — | default | Chat Completions |
o1-pro-2025-03-19 | Usage | $180 / 1M tokens | $720 / 1M tokens | $90 / 1M tokens | — | default | Chat Completions |
o1-2024-12-17 | Usage | $15 / 1M tokens | $60 / 1M tokens | $7.5 / 1M tokens | — | default | Chat Completions |
o1-mini-2024-09-12 | Usage | $1.1 / 1M tokens | $4.4 / 1M tokens | $0.55 / 1M tokens | — | default | Chat Completions |
o1-preview-2024-09-12 | Usage | $15 / 1M tokens | $60 / 1M tokens | $7.5 / 1M tokens | — | default | Chat Completions |
omni-moderation-latest | Usage | $0.2 / 1M tokens | $0.2 / 1M tokens | — | — | default | Chat Completions |
gpt-oss-120b | Usage | $0.5 / 1M tokens | $2 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
gpt-oss-20b | Usage | $0.1 / 1M tokens | $0.4 / 1M tokens | $0.01 / 1M tokens | — | default | Chat Completions |
sora-character | Per call | — | — | — | $0.01 / call | default | Chat Completions |
omni-moderation-2024-09-26 | Usage | $0.2 / 1M tokens | $0.2 / 1M tokens | — | — | default | Chat Completions |
Anthropic 10
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
claude-sonnet-5 | Usage | $2 / 1M tokens | $10 / 1M tokens | $0.2 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-8 | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-7 | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-7-thinking | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-6 | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-6-thinking | Usage | $5 / 1M tokens | $25 / 1M tokens | $0.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-sonnet-4-6 | Usage | $3 / 1M tokens | $15 / 1M tokens | $0.3 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-sonnet-4-6-thinking | Usage | $3 / 1M tokens | $15 / 1M tokens | $0.3 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-opus-4-20250514 | Usage | $15 / 1M tokens | $75 / 1M tokens | $1.5 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
claude-sonnet-4-20250514-thinking | Usage | $3 / 1M tokens | $15 / 1M tokens | $0.3 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
Google 19
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
gemini-3.7-flash | Usage | $0.75 / 1M tokens | $3.75 / 1M tokens | $0.075 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3.6-flash | Usage | $1.5 / 1M tokens | $7.5 / 1M tokens | $0.15 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3.5-flash | Usage | $1.5 / 1M tokens | $9 / 1M tokens | $0.15 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3.5-flash-lite | Usage | $0.3 / 1M tokens | $2.502 / 1M tokens | $0.03 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3.1-flash-image | Per call | — | — | — | $0.055 / call | default | Gemini, Chat Completions |
gemini-3.1-flash-image-preview | Per call | — | — | — | $0.055 / call | default | Gemini, Chat Completions |
gemini-3.1-flash-lite | Usage | $0.25 / 1M tokens | $1.5 / 1M tokens | $0.025 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3.1-flash-lite-image | Per call | — | — | — | $0.025 / call | default | Gemini, Chat Completions |
gemini-3.1-flash-lite-preview | Usage | $0.25 / 1M tokens | $1.5 / 1M tokens | $0.025 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3.1-pro-preview † | Usage | $2 / 1M tokens | $12 / 1M tokens | $0.2 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3-flash-preview | Usage | $0.44 / 1M tokens | $2.64 / 1M tokens | $0.044 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3-flash-preview-thinking | Usage | $0.44 / 1M tokens | $2.64 / 1M tokens | $0.044 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-3-pro-image | Per call | — | — | — | $0.09 / call | default | Gemini, Chat Completions |
gemini-3-pro-image-preview | Per call | — | — | — | $0.09 / call | default | Gemini, Chat Completions |
gemini-2.5-flash | Usage | $0.3 / 1M tokens | $2.4 / 1M tokens | $0.03 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-2.5-flash-image | Per call | — | — | — | $0.02 / call | default | Gemini, Chat Completions |
gemini-2.5-flash-lite | Usage | $0.1 / 1M tokens | $0.4 / 1M tokens | $0.01 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-2.5-flash-nothinking | Usage | $0.3 / 1M tokens | $2.4 / 1M tokens | $0.03 / 1M tokens | — | default | Gemini, Chat Completions |
gemini-2.5-pro | Usage | $1.25 / 1M tokens | $10 / 1M tokens | $0.125 / 1M tokens | — | default | Gemini, Chat Completions |
xAI 9
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
grok-4.6 † | Usage | $75 / 1M tokens | $75 / 1M tokens | $18.75 / 1M tokens | — | default | Chat Completions |
grok-4.5 † | Usage | $2 / 1M tokens | $6 / 1M tokens | $0.5 / 1M tokens | — | default | Chat Completions |
grok-4.3 † | Usage | $1.25 / 1M tokens | $2.5 / 1M tokens | $0.2 / 1M tokens | — | default | Chat Completions |
grok-4-1-fast-non-reasoning | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-1-fast-reasoning | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-fast-non-reasoning | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-4-fast-reasoning | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
grok-3 | Usage | $3 / 1M tokens | $15 / 1M tokens | — | — | default | Chat Completions |
grok-3-mini | Usage | $0.3 / 1M tokens | $1.8 / 1M tokens | — | — | default | Chat Completions |
DeepSeek 8
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
deepseek-v4-flash | Usage | $0.44 / 1M tokens | $1.32 / 1M tokens | $0.01408 / 1M tokens | — | default | Anthropic Messages, Chat Completions |
deepseek-v4-flash-vision-exp | Usage | $0.44 / 1M tokens | $1.32 / 1M tokens | $0.014 / 1M tokens | — | default | Anthropic Messages, Chat Completions |
deepseek-v4-pro | Usage | $1.74 / 1M tokens | $3.48 / 1M tokens | $0.057994 / 1M tokens | — | default | Anthropic Messages, Chat Completions |
DeepSeek-V3.2-Exp-nothinking | Usage | $0.3 / 1M tokens | $0.45 / 1M tokens | — | — | default | Chat Completions |
DeepSeek-V3.2-Exp-thinking | Usage | $0.3 / 1M tokens | $0.45 / 1M tokens | — | — | default | Chat Completions |
deepseek-v3.2 | Usage | $0.28 / 1M tokens | $0.42 / 1M tokens | — | — | default | Chat Completions |
deepseek-v3.2-thinking | Usage | $0.28 / 1M tokens | $0.42 / 1M tokens | — | — | default | Chat Completions |
deepseek-r1 | Usage | $0.57 / 1M tokens | $2.28 / 1M tokens | $0.1425 / 1M tokens | — | default | Chat Completions |
阿里巴巴 77
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
qwen3.6-flash † | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | default | Chat Completions |
qwen3.6-max-preview † | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | default | Chat Completions |
qwen3.6-plus † | Usage | $0.3 / 1M tokens | $1.8 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-122b-a10b | Usage | $0.12 / 1M tokens | $0.96 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-27b | Usage | $0.09 / 1M tokens | $0.72 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-35b-a3b | Usage | $0.06 / 1M tokens | $0.48 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-397b-a17b | Usage | $0.18 / 1M tokens | $1.08 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-flash † | Usage | $0.03 / 1M tokens | $0.3 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-plus † | Usage | $0.114 / 1M tokens | $0.684 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-flash-2026-02-23 † | Usage | $0.03 / 1M tokens | $0.3 / 1M tokens | — | — | default | Chat Completions |
qwen3.5-plus-2026-02-15 † | Usage | $0.114 / 1M tokens | $0.684 / 1M tokens | — | — | default | Chat Completions |
qwen3-235b-a22b | Usage | $1 / 1M tokens | $10 / 1M tokens | — | — | default | Chat Completions |
qwen3-30b-a3b | Usage | $0.2 / 1M tokens | $2 / 1M tokens | — | — | default | Chat Completions |
qwen3-32b | Usage | $0.4 / 1M tokens | $4 / 1M tokens | — | — | default | Chat Completions |
qwen3-coder-480b-a35b-instruct † | Usage | $3 / 1M tokens | $15 / 1M tokens | — | — | default | Chat Completions |
qwen3-coder-flash † | Usage | $0.5 / 1M tokens | $5 / 1M tokens | — | — | default | Chat Completions |
qwen3-coder-plus † | Usage | $5 / 1M tokens | $20 / 1M tokens | — | — | default | Chat Completions |
qwen3-max † | Usage | $1.2 / 1M tokens | $6 / 1M tokens | — | — | default | Chat Completions |
qwen3-max-preview † | Usage | $1.2 / 1M tokens | $6 / 1M tokens | — | — | default | Chat Completions |
qwen3-next-80b-a3b-instruct | Usage | $0.15 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
qwen3-omni-flash | Usage | $2 / 1M tokens | $20 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-235b-a22b-instruct | Usage | $0.3 / 1M tokens | $3 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-235b-a22b-thinking | Usage | $0.3 / 1M tokens | $3 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-30b-a3b-thinking | Usage | $0.12 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-32b-thinking | Usage | $0.3 / 1M tokens | $3 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-embedding | Usage | $0.25 / 1M tokens | $0.25 / 1M tokens | — | — | default | Embeddings, Chat Completions |
qwen3-vl-flash | Usage | $0.1 / 1M tokens | $0.8 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-plus † | Usage | $0.3 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-flash-2025-10-15 | Usage | $0.1 / 1M tokens | $0.8 / 1M tokens | — | — | default | Chat Completions |
qwen3-coder-plus-2025-09-23 † | Usage | $2 / 1M tokens | $20 / 1M tokens | — | — | default | Chat Completions |
qwen3-max-2025-09-23 † | Usage | $1.2 / 1M tokens | $6 / 1M tokens | — | — | default | Chat Completions |
qwen3-vl-plus-2025-09-23 † | Usage | $0.3 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
qwen3-omni-flash-2025-09-15 | Usage | $2 / 1M tokens | $20 / 1M tokens | — | — | default | Chat Completions |
qwen3-coder-plus-2025-07-22 † | Usage | $2 / 1M tokens | $20 / 1M tokens | — | — | default | Chat Completions |
qwen3-235b-a22b-instruct-2507 | Usage | $1 / 1M tokens | $10 / 1M tokens | — | — | default | Chat Completions |
qwen3-235b-a22b-thinking-2507 | Usage | $1.6 / 1M tokens | $12.8 / 1M tokens | — | — | default | Chat Completions |
qwen3-30b-a3b-instruct-2507 | Usage | $0.2 / 1M tokens | $0.8 / 1M tokens | — | — | default | Chat Completions |
qwen3-30b-a3b-thinking-2507 | Usage | $0.2 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
wan2.7-i2v | Usage | $1.2 / 1M tokens | $1.2 / 1M tokens | — | — | Wan | Chat Completions |
wan2.7-r2v | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | Wan | Chat Completions |
wan2.7-t2v | Usage | $1.2 / 1M tokens | $1.2 / 1M tokens | — | — | Wan | Chat Completions |
wan2.7-videoedit | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | Wan | Chat Completions |
wan2.6-i2v | Usage | $1.2 / 1M tokens | $1.2 / 1M tokens | — | — | Wan | Chat Completions |
wan2.6-r2v | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | Wan | Chat Completions |
wan2.6-r2v-flash | Usage | $75 / 1M tokens | $75 / 1M tokens | — | — | Wan | Chat Completions |
wan2.6-t2v | Usage | $1.2 / 1M tokens | $1.2 / 1M tokens | — | — | Wan | Chat Completions |
qwen2-72b-instruct | Usage | $2.8572 / 1M tokens | $2.8572 / 1M tokens | — | — | default | Chat Completions |
multimodal-embedding-v1 | Usage | $0 / 1M tokens | $0 / 1M tokens | — | — | default | Embeddings, Chat Completions |
qvq-max-latest | Usage | $1.2 / 1M tokens | $4.8 / 1M tokens | — | — | default | Chat Completions |
qwen-max-latest | Usage | $1.6 / 1M tokens | $6.4 / 1M tokens | — | — | default | Chat Completions |
qwen-plus-latest † | Usage | $0.4 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
qwen-turbo-latest | Usage | $0.2 / 1M tokens | $0.6 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-max-latest | Usage | $0.2 / 1M tokens | $0.8 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-ocr-latest | Usage | $0.044 / 1M tokens | $0.07348 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-plus-latest | Usage | $0.2 / 1M tokens | $0.6 / 1M tokens | — | — | default | Chat Completions |
qwq-plus-latest | Usage | $0.8 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
qvq-max | Usage | $1.2 / 1M tokens | $4.8 / 1M tokens | — | — | default | Chat Completions |
qvq-plus | Usage | $0.28 / 1M tokens | $0.7 / 1M tokens | — | — | default | Chat Completions |
qwen-long | Usage | $0.07 / 1M tokens | $0.28 / 1M tokens | — | — | default | Chat Completions |
qwen-max | Usage | $1.6 / 1M tokens | $6.4 / 1M tokens | — | — | default | Chat Completions |
qwen-max-longcontext | Usage | $3.2 / 1M tokens | $3.2 / 1M tokens | — | — | default | Chat Completions |
qwen-mt-plus | Usage | $2 / 1M tokens | $8 / 1M tokens | — | — | default | Chat Completions |
qwen-mt-turbo | Usage | $0.2 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
qwen-plus † | Usage | $0.4 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
qwen-turbo | Usage | $0.2 / 1M tokens | $0.6 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-max | Usage | $0.2 / 1M tokens | $0.8 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-ocr | Usage | $0.72 / 1M tokens | $0.72 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-plus | Usage | $0.2 / 1M tokens | $0.6 / 1M tokens | — | — | default | Chat Completions |
qwq-32b | Usage | $0.4 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
qwq-plus | Usage | $0.8 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
qwen-vl-ocr-2025-11-20 | Usage | $0.044 / 1M tokens | $0.07348 / 1M tokens | — | — | default | Chat Completions |
qwen-plus-2025-09-11 † | Usage | $0.4 / 1M tokens | $4 / 1M tokens | — | — | default | Chat Completions |
qwen-turbo-2025-07-15 | Usage | $0.2 / 1M tokens | $1.6 / 1M tokens | — | — | default | Chat Completions |
qwen-plus-2025-07-14 | Usage | $0.4 / 1M tokens | $4 / 1M tokens | — | — | default | Chat Completions |
qvq-max-2025-05-15 | Usage | $1 / 1M tokens | $4 / 1M tokens | — | — | default | Chat Completions |
qvq-plus-2025-05-15 | Usage | $0.28 / 1M tokens | $0.7 / 1M tokens | — | — | default | Chat Completions |
qwq-plus-2025-03-05 | Usage | $0.8 / 1M tokens | $2.4 / 1M tokens | — | — | default | Chat Completions |
字节跳动 10
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
seedream-5-0-260128 | Per call | — | — | — | $0.035 / call | default | Chat Completions |
seedream-4-5-251128 | Per call | — | — | — | $0.045 / call | default | Chat Completions |
seedream-4-0-250828 | Per call | — | — | — | $0.035 / call | default | Chat Completions |
doubao-seedance-2-0-mini-260615 | Usage | $23 / 1M tokens | $23 / 1M tokens | — | — | SeeDance2 | Chat Completions |
seed-2-0-lite-260428 | Usage | $0.25 / 1M tokens | $2 / 1M tokens | — | — | default | Chat Completions |
seed-2-0-mini-260428 | Usage | $0.1 / 1M tokens | $0.4 / 1M tokens | — | — | default | Chat Completions |
seed-2-0-code-preview-260328 | Usage | $0.5 / 1M tokens | $3 / 1M tokens | — | — | default | Chat Completions |
seed-2-0-pro-260328 † | Usage | $0.5 / 1M tokens | $3 / 1M tokens | — | — | default | Chat Completions |
doubao-seedance-2-0-260128 | Usage | $46 / 1M tokens | $46 / 1M tokens | — | — | SeeDance2 | Chat Completions |
doubao-seedance-2-0-fast-260128 | Usage | $37 / 1M tokens | $37 / 1M tokens | — | — | SeeDance2 | Chat Completions |
智谱 10
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
glm-5.2 | Usage | $1.142 / 1M tokens | $3.997 / 1M tokens | $0.2284 / 1M tokens | — | claude_code, default | Anthropic Messages, Chat Completions |
glm-5.1 † | Usage | $0.84 / 1M tokens | $3.36 / 1M tokens | $0.084 / 1M tokens | — | default | Chat Completions |
glm-5 † | Usage | $0.56 / 1M tokens | $2.52 / 1M tokens | $0.056 / 1M tokens | — | default | Chat Completions |
glm-4.7 | Usage | $0.6 / 1M tokens | $2.16 / 1M tokens | $0.06 / 1M tokens | — | default | Chat Completions |
glm-4.6 | Usage | $0.5 / 1M tokens | $2 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
glm-4.6v | Usage | $0.28 / 1M tokens | $0.84 / 1M tokens | $0.028 / 1M tokens | — | default | Chat Completions |
glm-4.5 | Usage | $0.5 / 1M tokens | $2 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
glm-4.5-air | Usage | $0.2 / 1M tokens | $1 / 1M tokens | $0.02 / 1M tokens | — | default | Chat Completions |
glm-4.5-flash | Usage | $0.01 / 1M tokens | $0.04 / 1M tokens | $0.001 / 1M tokens | — | default | Chat Completions |
glm-4.5v | Usage | $0.5 / 1M tokens | $1.5 / 1M tokens | $0.05 / 1M tokens | — | default | Chat Completions |
Moonshot 5
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
kimi-k2.6 | Usage | $0.95 / 1M tokens | $4 / 1M tokens | $0.16 / 1M tokens | — | default | Chat Completions |
kimi-k2.5 | Usage | $0.6 / 1M tokens | $3.15 / 1M tokens | $0.1 / 1M tokens | — | default | Chat Completions |
kimi-k2 | Usage | $0.56 / 1M tokens | $2.24 / 1M tokens | $0.056 / 1M tokens | — | default | Chat Completions |
kimi-k2-128k | Usage | $0.56 / 1M tokens | $2.24 / 1M tokens | $0.056 / 1M tokens | — | default | Chat Completions |
kimi-k2-instruct | Usage | $0.56 / 1M tokens | $0.56 / 1M tokens | $0.056 / 1M tokens | — | default | Chat Completions |
MiniMax 2
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
MiniMax-M2.5 | Usage | $0.3 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
MiniMax-M2.1 | Usage | $0.3 / 1M tokens | $1.2 / 1M tokens | — | — | default | Chat Completions |
Black Forest Labs 5
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
flux-2-flex | Per call | — | — | — | $0.06 / call | default | Images, Chat Completions |
flux-2-max | Per call | — | — | — | $0.07 / call | default | Images, Chat Completions |
flux-2-pro | Per call | — | — | — | $0.03 / call | default | Images, Chat Completions |
flux-kontext-max | Per call | — | — | — | $0.07 / call | default | Images, Chat Completions |
flux-kontext-pro | Per call | — | — | — | $0.035 / call | default | Images, Chat Completions |
美团 1
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
longcat-flash-chat | Usage | $0.25 / 1M tokens | $0.5 / 1M tokens | — | — | default | Chat Completions |
阶跃星辰 1
| Model ID | Billing | Input | Output | Cache read | Per-call | Token groups | Endpoints |
|---|---|---|---|---|---|---|---|
step-3.5-flash | Usage | $0.1 / 1M tokens | $0.3 / 1M tokens | — | — | default | Chat Completions |
Tiered pricing
Models marked with † use different price tiers based on the token count of a single request. The table below shows the input and output price for each tier. Confirm the final token count and charge in call logs.
| Model ID | Tokens per request | Input price | Output price |
|---|---|---|---|
qwen3.6-flash | 0–262,144 | $75 / 1M tokens | $450 / 1M tokens |
qwen3.6-flash | 262,145–1,024,000 | $300 / 1M tokens | $1800 / 1M tokens |
qwen3.6-max-preview | 0–131,072 | $75 / 1M tokens | $450 / 1M tokens |
qwen3.6-max-preview | 131,073–262,144 | $124.22 / 1M tokens | $745.31 / 1M tokens |
qwen3.6-plus | 0–262,144 | $0.3 / 1M tokens | $1.8 / 1M tokens |
qwen3.6-plus | 262,145–1,024,000 | $1.2 / 1M tokens | $7.2 / 1M tokens |
qwen3.5-flash | 0–131,072 | $0.03 / 1M tokens | $0.3 / 1M tokens |
qwen3.5-flash | 131,073–262,144 | $0.122143 / 1M tokens | $1.2214 / 1M tokens |
qwen3.5-flash | 262,145 and above | $0.182143 / 1M tokens | $1.8214 / 1M tokens |
qwen3.5-plus | 0–131,072 | $0.114 / 1M tokens | $0.684 / 1M tokens |
qwen3.5-plus | 131,073–262,144 | $0.28 / 1M tokens | $1.68 / 1M tokens |
qwen3.5-plus | 262,145 and above | $0.56 / 1M tokens | $3.36 / 1M tokens |
qwen3.5-flash-2026-02-23 | 0–131,072 | $0.03 / 1M tokens | $0.3 / 1M tokens |
qwen3.5-flash-2026-02-23 | 131,073–262,144 | $0.122143 / 1M tokens | $1.2214 / 1M tokens |
qwen3.5-flash-2026-02-23 | 262,145 and above | $0.182143 / 1M tokens | $1.8214 / 1M tokens |
qwen3.5-plus-2026-02-15 | 0–131,072 | $0.114 / 1M tokens | $0.684 / 1M tokens |
qwen3.5-plus-2026-02-15 | 131,073–262,144 | $0.28 / 1M tokens | $1.68 / 1M tokens |
qwen3.5-plus-2026-02-15 | 262,145 and above | $0.56 / 1M tokens | $3.36 / 1M tokens |
qwen3-coder-480b-a35b-instruct | 0–32,768 | $3 / 1M tokens | $15 / 1M tokens |
qwen3-coder-480b-a35b-instruct | 32,769–131,072 | $5.4 / 1M tokens | $27 / 1M tokens |
qwen3-coder-480b-a35b-instruct | 131,073–204,800 | $9 / 1M tokens | $45 / 1M tokens |
qwen3-coder-flash — first pricing tier; confirm the current console price | 0–32,000 | $0.5 / 1M tokens | $2 / 1M tokens |
qwen3-coder-flash | 32,001–128,000 | $0.75 / 1M tokens | $3 / 1M tokens |
qwen3-coder-flash | 128,001–256,000 | $1.25 / 1M tokens | $5 / 1M tokens |
qwen3-coder-flash | 256,001–1,000,000 | $2.5 / 1M tokens | $12.5 / 1M tokens |
qwen3-coder-plus | 0–32,000 | $5 / 1M tokens | $20 / 1M tokens |
qwen3-coder-plus | 32,001–128,000 | $7.5 / 1M tokens | $30 / 1M tokens |
qwen3-coder-plus | 128,001–256,000 | $12.5 / 1M tokens | $50 / 1M tokens |
qwen3-coder-plus | 256,001–1,000,000 | $25 / 1M tokens | $250 / 1M tokens |
qwen3-max — first pricing tier; confirm the current console price | 0–32,000 | $1.2 / 1M tokens | $4.8 / 1M tokens |
qwen3-max | 32,001–128,000 | $1.92 / 1M tokens | $7.68 / 1M tokens |
qwen3-max | 128,001–256,000 | $3.36 / 1M tokens | $13.44 / 1M tokens |
qwen3-max-preview — first pricing tier; confirm the current console price | 0–32,000 | $1.2 / 1M tokens | $4.8 / 1M tokens |
qwen3-max-preview | 32,001–128,000 | $2 / 1M tokens | $8 / 1M tokens |
qwen3-max-preview | 128,001–256,000 | $3 / 1M tokens | $12 / 1M tokens |
qwen3-vl-plus — first pricing tier; confirm the current console price | 0–32,000 | $0.3 / 1M tokens | $3 / 1M tokens |
qwen3-vl-plus | 32,001–128,000 | $0.45 / 1M tokens | $4.5 / 1M tokens |
qwen3-vl-plus | 128,001–256,000 | $0.9 / 1M tokens | $9 / 1M tokens |
qwen3-coder-plus-2025-09-23 — first pricing tier; confirm the current console price | 0–32,000 | $2 / 1M tokens | $8 / 1M tokens |
qwen3-coder-plus-2025-09-23 | 32,001–128,000 | $3 / 1M tokens | $12 / 1M tokens |
qwen3-coder-plus-2025-09-23 | 128,001–256,000 | $5 / 1M tokens | $20 / 1M tokens |
qwen3-coder-plus-2025-09-23 | 256,001–1,000,000 | $10 / 1M tokens | $100 / 1M tokens |
qwen3-max-2025-09-23 — first pricing tier; confirm the current console price | 0–32,000 | $1.2 / 1M tokens | $4.8 / 1M tokens |
qwen3-max-2025-09-23 | 32,001–128,000 | $2 / 1M tokens | $8 / 1M tokens |
qwen3-max-2025-09-23 | 128,001–256,000 | $3 / 1M tokens | $12 / 1M tokens |
qwen3-vl-plus-2025-09-23 — first pricing tier; confirm the current console price | 0–32,000 | $0.3 / 1M tokens | $3 / 1M tokens |
qwen3-vl-plus-2025-09-23 | 32,001–128,000 | $0.45 / 1M tokens | $4.5 / 1M tokens |
qwen3-vl-plus-2025-09-23 | 128,001–256,000 | $0.9 / 1M tokens | $9 / 1M tokens |
qwen3-coder-plus-2025-07-22 — first pricing tier; confirm the current console price | 0–32,000 | $2 / 1M tokens | $8 / 1M tokens |
qwen3-coder-plus-2025-07-22 | 32,001–128,000 | $3 / 1M tokens | $12 / 1M tokens |
qwen3-coder-plus-2025-07-22 | 128,001–256,000 | $5 / 1M tokens | $20 / 1M tokens |
qwen3-coder-plus-2025-07-22 | 256,001–1,000,000 | $10 / 1M tokens | $100 / 1M tokens |
qwen-plus-latest — first pricing tier; confirm the current console price | 0–128,000 | $0.4 / 1M tokens | $1 / 1M tokens |
qwen-plus-latest | 128,001–256,000 | $1.2 / 1M tokens | $10 / 1M tokens |
qwen-plus-latest | 256,001–1,000,000 | $2.4 / 1M tokens | $24 / 1M tokens |
qwen-plus — first pricing tier; confirm the current console price | 0–128,000 | $0.4 / 1M tokens | $1 / 1M tokens |
qwen-plus | 128,001–256,000 | $1.2 / 1M tokens | $10 / 1M tokens |
qwen-plus | 256,001–1,000,000 | $2.4 / 1M tokens | $24 / 1M tokens |
qwen-plus-2025-09-11 — first pricing tier; confirm the current console price | 0–128,000 | $0.4 / 1M tokens | $1 / 1M tokens |
qwen-plus-2025-09-11 | 128,001–256,000 | $1.2 / 1M tokens | $10 / 1M tokens |
qwen-plus-2025-09-11 | 256,001–1,000,000 | $2.4 / 1M tokens | $24 / 1M tokens |
glm-5.1 | 0–32,768 | $0.84 / 1M tokens | $3.36 / 1M tokens |
glm-5.1 | 32,769 and above | $1.14 / 1M tokens | $3.99 / 1M tokens |
glm-5 | 0–32,000 | $0.56 / 1M tokens | $2.52 / 1M tokens |
glm-5 | 32,001 and above | $0.86 / 1M tokens | $3.096 / 1M tokens |
seed-2-0-pro-260328 | 0–131,072 | $0.5 / 1M tokens | $3 / 1M tokens |
seed-2-0-pro-260328 | 131,073–262,144 | $1 / 1M tokens | $6 / 1M tokens |
gemini-3.1-pro-preview | 0–200,000 | $2 / 1M tokens | $12 / 1M tokens |
gemini-3.1-pro-preview | 200,001 and above | $4 / 1M tokens | $18 / 1M tokens |
gpt-5.6-luna | 0–272,000 | $0.2 / 1M tokens | $1.2 / 1M tokens |
gpt-5.6-luna | 272,001 and above | $0.4 / 1M tokens | $1.8 / 1M tokens |
gpt-5.6-sol | 0–272,000 | $5 / 1M tokens | $30 / 1M tokens |
gpt-5.6-sol | 272,001 and above | $10 / 1M tokens | $45 / 1M tokens |
gpt-5.6-terra | 0–272,000 | $2 / 1M tokens | $12 / 1M tokens |
gpt-5.6-terra | 272,001 and above | $4 / 1M tokens | $18 / 1M tokens |
gpt-5.5 | 0–278,528 | $5 / 1M tokens | $30 / 1M tokens |
gpt-5.5 | 278,529 and above | $10 / 1M tokens | $45 / 1M tokens |
gpt-5.4 | 0–278,528 | $2.5 / 1M tokens | $15 / 1M tokens |
gpt-5.4 | 278,529 and above | $5 / 1M tokens | $22.5 / 1M tokens |
gpt-5.4-pro | 0–278,528 | $30 / 1M tokens | $180 / 1M tokens |
gpt-5.4-pro | 278,529 and above | $60 / 1M tokens | $270 / 1M tokens |
grok-4.6 | 0–204,800 | $75 / 1M tokens | $225 / 1M tokens |
grok-4.6 | 204,801–512,000 | $150 / 1M tokens | $450 / 1M tokens |
grok-4.5 | 0–204,800 | $2 / 1M tokens | $6 / 1M tokens |
grok-4.5 | 204,801 and above | $4 / 1M tokens | $12 / 1M tokens |
grok-4.3 | 0–204,800 | $1.25 / 1M tokens | $2.5 / 1M tokens |
grok-4.3 | 204,801 and above | $2.5 / 1M tokens | $5 / 1M tokens |
Frequently asked questions
Is this price read in real time?
This page is generated periodically from current pricing configuration and shows its update time. Model changes, account groups, or dedicated contracts can produce a different console price; confirm final charges in the console and call logs.
When does the cache-read price apply?
Only when the request hits a supported prompt cache. If the model has no cache setting, the cache is missed, or the calling method does not support caching, normal input pricing applies.
How does tiered pricing work?
The system selects a tier from the token count of each request. The tier table lists both input and output prices; confirm final usage and charges in call logs.
How do I confirm enterprise purchasing, volume usage, or discounts?
Contact the site owner or support with the account, models, expected usage, contract, and invoice requirements. The documentation does not promise a fixed discount threshold or percentage.
How do I call a model after finding its ID?
Confirm the endpoint and token group shown in the table, then create the matching token in the console. OpenAI-compatible models generally use /v1/chat/completions or /v1/responses; Gemini, image, and specialized models should follow their dedicated documentation.
Pricing snapshot generated: 2026-09-02T04:12:17.377Z