Pricing
Simple per-token pricing. No minimums, no commitments. All prices in USD per 1M tokens.
| Model | Context | Max Output | Input Price | Output Price | Cache Read |
|---|---|---|---|---|---|
glm-5.3 1.3M context flagship reasoning model | 1.3M | 65.5K | $0.756 / 1M tokens | $2.38 / 1M tokens | $0.140 / 1M tokens |
glm-5.3-flash Fast, cost-efficient 1.3M context model | 1.3M | 65.5K | $0.135 / 1M tokens | $0.450 / 1M tokens | $0.045 / 1M tokens |
glm-5.2 Flagship reasoning model | 203K | 65.5K | $1.26 / 1M tokens | $3.96 / 1M tokens | $0.234 / 1M tokens |
glm-5.1 Flagship reasoning model | 203K | 65.5K | $0.931 / 1M tokens | $2.93 / 1M tokens | $0.173 / 1M tokens |
glm-5 203K context reasoning model | 203K | 65.5K | $0.570 / 1M tokens | $1.82 / 1M tokens | $0.114 / 1M tokens |
glm-4.7 Flagship reasoning model | 200K | 65.5K | $0.380 / 1M tokens | $1.66 / 1M tokens | $0.076 / 1M tokens |
glm-4.7-flash Fast, cost-efficient variant | 200K | 65.5K | $0.057 / 1M tokens | $0.380 / 1M tokens | $0.0095 / 1M tokens |
glm-4.6 Agentic coding and reasoning | 200K | 65.5K | $0.380 / 1M tokens | $1.55 / 1M tokens | $0.070 / 1M tokens |
glm-4.5 General-purpose model | 131K | 65.5K | $0.570 / 1M tokens | $2.09 / 1M tokens | $0.105 / 1M tokens |
glm-4.5-air Lightweight, budget-friendly | 131K | 65.5K | $0.123 / 1M tokens | $0.807 / 1M tokens | $0.024 / 1M tokens |
kimi-k3 1M context reasoning model | 1M | 65.5K | $2.70 / 1M tokens | $13.50 / 1M tokens | $0.270 / 1M tokens |
kimi-k2.7-code Coding-focused reasoning model | 262K | 65.5K | $0.635 / 1M tokens | $2.97 / 1M tokens | $0.162 / 1M tokens |
kimi-k2.6 262K context reasoning model | 262K | 65.5K | $0.855 / 1M tokens | $3.60 / 1M tokens | $0.144 / 1M tokens |
kimi-k2.5 262K context, MoE architecture | 262K | 65.5K | $0.356 / 1M tokens | $1.98 / 1M tokens | $0.225 / 1M tokens |
minimax-m3 1M context, always-on reasoning | 1M | 65.5K | $0.270 / 1M tokens | $1.08 / 1M tokens | $0.054 / 1M tokens |
minimax-m2.7 204.8K context, always-on reasoning | 204.8K | 65.5K | $0.270 / 1M tokens | $1.08 / 1M tokens | $0.054 / 1M tokens |
minimax-m2.5 196.6K context, always-on reasoning | 196.6K | 65.5K | $0.143 / 1M tokens | $0.855 / 1M tokens | $0.040 / 1M tokens |
deepseek-v4-pro 1M context flagship reasoning model | 1M | 65.5K | $0.851 / 1M tokens | $1.70 / 1M tokens | $0.071 / 1M tokens |
deepseek-v4-flash Fast, cost-efficient 1M context model | 1M | 65.5K | $0.065 / 1M tokens | $0.129 / 1M tokens | $0.013 / 1M tokens |
qwen3-coder-next 262K context, fast code generation | 262K | 65.5K | $0.108 / 1M tokens | $0.675 / 1M tokens | $0.060 / 1M tokens |
Pricing is subject to change. All prices are in USD per 1M tokens. Most models support text input/output, tool calling, JSON mode, and streaming. Reasoning is supported on every model except Qwen3 Coder Next. MiniMax models always reason.