Skip to content

Models & pricing

Nine production models, priced in USD per 1M tokens. Upstream IDs public.

ModelCapabilitiesContextInputCached inputOutput
glm-5.240% off@cf/zai-org/glm-5.2toolsreasoningjson262K$1.40$0.84$0.26$0.156$4.40$2.64
kimi-k2.7-code40% off@cf/moonshotai/kimi-k2.7-codetoolsvisionreasoning262K$0.95$0.57$0.19$0.114$4.00$2.40
kimi-k2.640% off@cf/moonshotai/kimi-k2.6toolsvisionreasoning262K$0.95$0.57$0.16$0.096$4.00$2.40
glm-4.7-flash40% off@cf/zai-org/glm-4.7-flashtoolsreasoningfast131K$0.06$0.036—$0.40$0.24
deepseek-v4-flash40% off@cf/deepseek-ai/deepseek-v4-flash-0731toolsreasoningfast1M$0.44$0.264$0.014$0.008$1.32$0.792
deepseek-v4-pro40% off@cf/deepseek-ai/deepseek-v4-pro-0813toolsreasoningjson1M$1.32$0.792$0.044$0.026$3.96$2.376
qwen3.8-27b40% off@cf/qwen/qwen3.8-27btoolsreasoningvision262K$0.45$0.27$0.05$0.03$3.20$1.92
glm-5.3-flash40% off@cf/zai-org/glm-5.3-flashtoolsvisionreasoningjson1M$0.15$0.09$0.03$0.018$0.50$0.30
glm-5.340% off@cf/zai-org/glm-5.3toolsreasoningjson1M$1.40$0.84$0.26$0.156$4.40$2.64

Full cache hits are billed at 10% of the normal price. Prices as of 2026-09-25 — dashboard shows live rates.

glm-5.2

40% off@cf/zai-org/glm-5.2

Z.ai’s proven reasoning workhorse. Strong at multi-step reasoning, math and long-context analysis, with tool calls and strict JSON output.

Context window262K
Capabilitiestoolsreasoningjson

Reasoning tokens are billed at the output rate — never twice.

Input$1.40$0.84
Cached input$0.26$0.156
Output$4.40$2.64

kimi-k2.7-code

40% off@cf/moonshotai/kimi-k2.7-code

Moonshot AI's coding specialist. Tuned for repository-scale code work and agentic tool use, with vision input for screenshots and diagrams.

Context window262K
Capabilitiestoolsvisionreasoning

Reasoning tokens are billed at the output rate — never twice.

Input$0.95$0.57
Cached input$0.19$0.114
Output$4.00$2.40

kimi-k2.6

40% off@cf/moonshotai/kimi-k2.6

General-purpose Kimi model balancing quality and cost. Strong tool calling and vision — a solid default for assistants and agents.

Context window262K
Capabilitiestoolsvisionreasoning

Reasoning tokens are billed at the output rate — never twice.

Input$0.95$0.57
Cached input$0.16$0.096
Output$4.00$2.40

glm-4.7-flash

40% off@cf/zai-org/glm-4.7-flash

The fast, low-cost workhorse. Median time to first token is about 0.65 s in our own measurements — built for high-volume classification, extraction and chat.

Context window131K
Capabilitiestoolsreasoningfast

Reasoning tokens are billed at the output rate — never twice.

Input$0.06$0.036
Cached input—
Output$0.40$0.24

deepseek-v4-flash

40% off@cf/deepseek-ai/deepseek-v4-flash-0731

DeepSeek's official V4 Flash release with a 1M-token context window. Fast in our tests (~1 s to first token, ~100 tok/s) with strong agentic tool calling — the value pick for long-context work.

Context window1M
Capabilitiestoolsreasoningfast

Reasoning tokens are billed at the output rate — never twice.

Input$0.44$0.264
Cached input$0.014$0.008
Output$1.32$0.792

deepseek-v4-pro

40% off@cf/deepseek-ai/deepseek-v4-pro-0813

DeepSeek's flagship V4 model: 1M-token context window, deep reasoning and reliable function calling. Built for repository-scale analysis and long-document workloads.

Context window1M
Capabilitiestoolsreasoningjson

Reasoning tokens are billed at the output rate — never twice.

Input$1.32$0.792
Cached input$0.044$0.026
Output$3.96$2.376

qwen3.8-27b

40% off@cf/qwen/qwen3.8-27b

Alibaba's Qwen 3.8 27B — image input verified by our own probe, not copied from a spec sheet. Send images alongside text in the same request, with a 262K-token context window, tool calling and reasoning.

Context window262K
Capabilitiestoolsreasoningvision

Reasoning tokens are billed at the output rate — never twice.

Input$0.45$0.27
Cached input$0.05$0.03
Output$3.20$1.92

glm-5.3-flash

40% off@cf/zai-org/glm-5.3-flash

Z.ai's GLM 5.3 Flash: a 1M-token context window and image input verified by our own probe, in one model — at the lowest input price of any vision model in the catalog. Tool calls, reasoning and strict JSON included.

Context window1M
Capabilitiestoolsvisionreasoningjson

Reasoning tokens are billed at the output rate — never twice.

Input$0.15$0.09
Cached input$0.03$0.018
Output$0.50$0.30

glm-5.3

40% off@cf/zai-org/glm-5.3

Z.ai’s flagship agentic coding model: a 1M-token context window at the same list price as GLM-5.2, with tool calls, five reasoning-effort levels and strict JSON output. Text-only — image input is refused up front, at no cost to you.

Context window1M
Capabilitiestoolsreasoningjson

Reasoning tokens are billed at the output rate — never twice.

Input$1.40$0.84
Cached input$0.26$0.156
Output$4.40$2.64

About reasoning tokens: reasoning models stream their thinking as part of the completion output. Those tokens are a subset of output tokens, billed once at the output rate, and shown separately in your usage logs — they are never billed on top.