Supported Models
179 models from 21 model makers. Call any model by its ID in the maker/model format.
Token prices are per 1M tokens; image/video prices are per image/second. All include our service fee.
92 models
OpenAI (33)
| Model ID | Capabilities | Context | Input | Output |
|---|---|---|---|---|
openai/gpt-3.5-turboOpenAI: GPT-3.5 Turbo The model that made cheap chat ordinary. Long since beaten on price and quality by small GPT-5 tiers; kept for legacy integrations. | Tools | 16K | $0.600 | $1.80 |
openai/gpt-3.5-turbo-16kOpenAI: GPT-3.5 Turbo 16k The extended-context GPT-3.5 variant, now priced above its own successors. Legacy only. | Tools | 16K | $3.60 | $4.80 |
openai/gpt-4OpenAI: GPT-4 The original GPT-4, with an 8K window and the highest token price in the catalog. Present for legacy compatibility, not for new work. | Tools | 8K | $36.00 | $72.00 |
openai/gpt-4-turboOpenAI: GPT-4 Turbo The 128K GPT-4 generation that preceded 4o. Superseded on both price and speed, kept for older integrations. | VisionTools | 128K | $12.00 | $36.00 |
openai/gpt-4.1OpenAI: GPT-4.1 A million-token GPT-4 generation with vision and tool calling, and a long-standing default for document-heavy work. Later GPT-5 tiers usually beat it on price for the same job. | VisionTools | 1.0M | $2.40 | $9.60 |
openai/gpt-4.1-miniOpenAI: GPT-4.1 Mini The mid-size GPT-4.1, keeping the million-token window at a fifth of the output price. A practical choice for long-context work on a budget. | VisionTools | 1.0M | $0.480 | $1.92 |
openai/gpt-4.1-nanoOpenAI: GPT-4.1 Nano The smallest GPT-4.1, and one of the cheapest models here that still reads a million tokens. Suited to bulk summarisation and extraction over large corpora. | VisionTools | 1.0M | $0.120 | $0.480 |
openai/gpt-4oOpenAI: GPT-4o OpenAI's multimodal workhorse from the 4-series: text, vision and tool calling over a 128K window. Widely integrated and well understood, though GPT-5 tiers are cheaper for most new work. | VisionTools | 128K | $3.00 | $12.00 |
openai/gpt-4o-2024-05-13OpenAI: GPT-4o (2024-05-13) The first GPT-4o snapshot, priced above later ones. Only worth calling if you are reproducing results from that exact version. | VisionTools | 128K | $6.00 | $18.00 |
openai/gpt-4o-2024-08-06OpenAI: GPT-4o (2024-08-06) A pinned GPT-4o snapshot, the one that introduced structured outputs. Kept for integrations that were validated against it. | VisionTools | 128K | $3.00 | $12.00 |
openai/gpt-4o-2024-11-20OpenAI: GPT-4o (2024-11-20) A pinned GPT-4o snapshot. Use it when you need output that cannot shift as OpenAI updates the rolling alias; otherwise call gpt-4o. | VisionTools | 128K | $3.00 | $12.00 |
openai/gpt-4o-miniOpenAI: GPT-4o-mini The small GPT-4o, priced for volume while keeping vision and tool calling. A common default for assistants that must stay cheap per turn. | VisionTools | 128K | $0.180 | $0.720 |
openai/gpt-4o-mini-2024-07-18OpenAI: GPT-4o-mini (2024-07-18) A pinned GPT-4o-mini snapshot, for reproducible output from the small 4o tier. | VisionTools | 128K | $0.180 | $0.720 |
openai/gpt-4o-mini-search-previewOpenAI: GPT-4o-mini Search Preview The small search-enabled GPT-4o, for grounding high-volume answers in current web results without paying full 4o prices. | Chat | 128K | $0.180 | $0.720 |
openai/gpt-4o-search-previewOpenAI: GPT-4o Search Preview GPT-4o wired to OpenAI's web search, for answers that need current information rather than training-cutoff knowledge. No tool calling: search is the tool. | Chat | 128K | $3.00 | $12.00 |
openai/gpt-5OpenAI: GPT-5 The first GPT-5 generation, with a 400K context window, vision and tool calling. Still a solid general-purpose default, though later 5.x tiers are usually the better buy. | VisionTools | 400K | $1.50 | $12.00 |
openai/gpt-5-miniOpenAI: GPT-5 Mini The mid-size GPT-5, five times cheaper than the full model on output. The usual pick for chat and tool loops that run at volume. | VisionTools | 400K | $0.300 | $2.40 |
openai/gpt-5-nanoOpenAI: GPT-5 Nano The cheapest GPT-5 tier by a wide margin. Built for classification, routing and short structured extraction at scale. | VisionTools | 400K | $0.060 | $0.480 |
openai/gpt-5.1OpenAI: GPT-5.1 An incremental GPT-5 release at the same price as GPT-5, with a 400K window. Useful as a pinned target when you need output that does not shift under you. | VisionTools | 400K | $1.50 | $12.00 |
openai/gpt-5.2OpenAI: GPT-5.2 A GPT-5 generation with a 400K context window, positioned between 5.1 and 5.4. Kept for integrations pinned to it. | VisionTools | 400K | $2.10 | $16.80 |
openai/gpt-5.4OpenAI: GPT-5.4 A million-token GPT-5 generation at half the output price of 5.5. Worth benchmarking against 5.5 before paying the difference. | VisionTools | 1.1M | $3.00 | $18.00 |
openai/gpt-5.4-miniOpenAI: GPT-5.4 Mini The mid-size GPT-5.4, six times cheaper than the full model on output. Sized for assistant traffic and tool-calling loops that run constantly. | VisionTools | 400K | $0.900 | $5.40 |
openai/gpt-5.4-nanoOpenAI: GPT-5.4 Nano The smallest GPT-5.4, priced for very high volume: routing, tagging, short extraction, and anything you call thousands of times an hour. | VisionTools | 400K | $0.240 | $1.50 |
openai/gpt-5.5OpenAI: GPT-5.5 OpenAI's broad multimodal frontier model with a million-token context. A generalist pick when a workload mixes long documents, images and tool use rather than specialising in one. | VisionTools | 1.1M | $6.00 | $36.00 |
openai/gpt-5.6-lunaOpenAI: GPT-5.6 Luna The economical GPT-5.6 tier, five times cheaper than Sol on output while keeping the million-token window. Sized for high-volume calls that still need frontier-family behaviour. | VisionTools | 1.1M | $1.20 | $7.20 |
openai/gpt-5.6-solOpenAI: GPT-5.6 Sol The top tier of the GPT-5.6 family, priced for work where answer quality decides the outcome rather than throughput. Million-token context with vision and tool calling. | VisionTools | 1.1M | $6.00 | $36.00 |
openai/gpt-5.6-terraOpenAI: GPT-5.6 Terra The middle GPT-5.6 tier, at half the price of Sol and twice that of Luna. The tier to try first when you want 5.6 behaviour on production traffic. | VisionTools | 1.1M | $3.00 | $18.00 |
openai/gpt-oss-120bOpenAI: GPT-OSS 120B OpenAI's large open-weight model, served here so you can call it without hosting it. Tool calling included; no vision. | Tools | 128K | $0.180 | $0.720 |
openai/gpt-oss-20bOpenAI: GPT-OSS 20B The small open-weight OpenAI model and one of the cheapest tool-calling options in the catalog. Good fit for agent loops where per-step cost dominates. | Tools | 128K | $0.084 | $0.360 |
openai/o1OpenAI: o1 OpenAI's first reasoning model, which thinks before answering rather than streaming an immediate reply. Later o-series releases deliver the same approach far more cheaply. | VisionTools | 200K | $18.00 | $72.00 |
openai/o3OpenAI: o3 A reasoning model for problems where a considered answer beats a fast one: maths, analysis, multi-step planning. Roughly a seventh of o1's output price. | VisionTools | 200K | $2.40 | $9.60 |
openai/o3-miniOpenAI: o3 Mini The small o3, for reasoning-shaped work that has to run at volume. Cheaper per call than most frontier chat models. | VisionTools | 200K | $1.32 | $5.28 |
openai/o4-miniOpenAI: o4 Mini A compact reasoning model at the same price as o3-mini, worth A/B testing against it on your own tasks before committing. | VisionTools | 200K | $1.32 | $5.28 |
Google (15)
| Model ID | Capabilities | Context | Input | Output |
|---|---|---|---|---|
google/gemini-2.5-flashGoogle: Gemini 2.5 Flash The Gemini 2.5 workhorse: a million tokens of context at a tenth of Pro's price. Still one of the better value picks for long-context summarisation. | VisionTools | 1.0M | $0.360 | $3.00 |
google/gemini-2.5-flash-imageGoogle: Nano Banana (Gemini 2.5 Flash Image) Gemini 2.5 Flash with image understanding, known as Nano Banana. A chat model that reads images rather than an image generator, on a 32K window. | Vision | 33K | $0.360 | $3.00 |
google/gemini-2.5-flash-liteGoogle: Gemini 2.5 Flash Lite The cheapest Gemini 2.5 tier, and among the lowest token prices in the whole catalog. Sized for very high-volume, low-complexity calls. | VisionTools | 1.0M | $0.120 | $0.480 |
google/gemini-2.5-proGoogle: Gemini 2.5 Pro The Gemini 2.5 flagship, with a million-token window, vision and tool calling. A strong long-document model, now undercut on price by the Gemini 3 line. | VisionTools | 1.0M | $1.50 | $12.00 |
google/gemini-3-flash-previewGoogle: Gemini 3 Flash Preview The Gemini 3 Flash preview, a million-token multimodal tier priced between 2.5 Flash and 3.5 Flash. | VisionTools | 1.0M | $0.600 | $3.60 |
google/gemini-3.1-flash-liteGoogle: Gemini 3.1 Flash Lite The cheapest Gemini 3.1 tier that still reads a million tokens. Built for bulk classification and summarisation where per-call cost decides the architecture. | VisionTools | 1.0M | $0.300 | $1.80 |
google/gemini-3.1-flash-lite-previewGoogle: Gemini 3.1 Flash Lite Preview The preview channel for Gemini 3.1 Flash Lite. Same shape as the stable tier; use it to test upcoming behaviour, not to serve production traffic. | VisionTools | 1.0M | $0.300 | $1.80 |
google/gemini-3.1-pro-previewGoogle: Gemini 3.1 Pro Preview The Pro tier of Gemini 3.1, for reasoning-heavy work over very long inputs. Preview, so expect behaviour to move before general availability. | VisionTools | 1.0M | $2.40 | $14.40 |
google/gemini-3.1-pro-preview-customtoolsGoogle: Gemini 3.1 Pro Preview Custom Tools Gemini 3.1 Pro with custom tool definitions enabled, for agents that call your own functions rather than Google's built-ins. | VisionTools | 1.0M | $2.40 | $14.40 |
google/gemini-3.5-flashGoogle: Gemini 3.5 Flash The current Flash generation: a million-token window, vision and tool calling at a fraction of Pro pricing. Google's usual answer for high-volume multimodal work. | VisionTools | 1.0M | $1.80 | $10.80 |
google/gemini-3.5-flash-liteGemini 3.5 Flash Lite Gemini 3.5 Flash Lite is Flash-Lite's first step into the agentic space, prioritizing thinking and tool calling to serve as a quick, efficient, and capable subagent for larger, more complex workflows. Compared to Gemini 3.1 Flash-Lite, it delivers stronger agentic performance and tool use, more precise and complete document understanding, | Chat | — | $0.345 | $2.88 |
google/gemini-3.6-flashGoogle: Gemini 3.6 Flash Gemini 3.6 Flash is designed to deliver strong coding and general-agentic capabilities (near-Pro level) at substantial speed and value with improved token-efficiency and quality over previous Flash models. Gemini 3.6 Flash is optimized for multi-step orchestration, full-stack code refactoring, and general reasoning with much better token efficiency than Gemini 3.5 Flash. | VisionTools | 1.0M | $1.73 | $8.63 |
google/gemini-3.7-flashGoogle: Gemini 3.7 Flash Gemini 3.7 Flash is Google's newest Flash generation model, aimed at coding and agentic workflows, with a 1M token context window and up to 64k output tokens. On ApexApi it runs as a text in, text out model at the same published price as Gemini 3.6 Flash, so it is the default pick when you want the newer model at that price point. Choose google/gemini-3.6-flash instead when your request needs image input or function calling. | Chat | 1.0M | $1.73 | $8.63 |
google/gemma-4-26b-a4b-itGoogle: Gemma 4 26B A4B The sparse 26B Gemma 4, which activates a fraction of its parameters per token and prices accordingly. One of the cheapest vision-capable models here. | VisionTools | 262K | $0.072 | $0.396 |
google/gemma-4-31b-itGoogle: Gemma 4 31B Google's open-weight instruction-tuned model at 31B, served so you do not have to host it. Multimodal input and tool calling at open-weight pricing. | VisionTools | 262K | $0.144 | $0.444 |
Mistral (14)
| Model ID | Capabilities | Context | Input | Output |
|---|---|---|---|---|
mistralai/codestral-2508Mistral: Codestral 2508 Mistral's code model, tuned for completion, refactoring and review rather than open-ended chat. 256K of context for repository-scale work. | Tools | 256K | $0.360 | $1.08 |
mistralai/devstral-2512Mistral: Devstral 2 2512 Mistral's agentic coding model, aimed at multi-step development tasks rather than single completions. | Tools | 262K | $0.480 | $2.40 |
mistralai/devstral-2512Mistral: Devstral 2 2512 Mistral's agentic coding model, aimed at multi-step development tasks rather than single completions. | Tools | 262K | $0.480 | $2.40 |
mistralai/ministral-14b-2512Mistral: Ministral 3 14B 2512 The largest Ministral, with symmetric input and output pricing that makes cost easy to predict for generation-heavy work. | VisionTools | 262K | $0.240 | $0.240 |
mistralai/ministral-14b-2512Mistral: Ministral 3 14B 2512 The largest Ministral, with symmetric input and output pricing that makes cost easy to predict for generation-heavy work. | VisionTools | 262K | $0.240 | $0.240 |
mistralai/ministral-3b-2512Mistral: Ministral 3 3B 2512 The smallest Ministral, and one of the cheapest models here. Built for edge-shaped workloads: routing, tagging, short replies. | VisionTools | 131K | $0.120 | $0.120 |
mistralai/ministral-3b-2512Mistral: Ministral 3 3B 2512 The smallest Ministral, and one of the cheapest models here. Built for edge-shaped workloads: routing, tagging, short replies. | VisionTools | 131K | $0.120 | $0.120 |
mistralai/ministral-8b-2512Mistral: Ministral 3 8B 2512 The mid Ministral at 8B, multimodal with tool calling, priced flat on input and output. | VisionTools | 262K | $0.180 | $0.180 |
mistralai/ministral-8b-2512Mistral: Ministral 3 8B 2512 The mid Ministral at 8B, multimodal with tool calling, priced flat on input and output. | VisionTools | 262K | $0.180 | $0.180 |
mistralai/mistral-large-2512Mistral: Mistral Large 3 2512 Mistral's flagship, with a 256K window, vision and tool calling, at a price well under most frontier models. A common European-hosted choice for production reasoning. | VisionTools | 262K | $0.600 | $1.80 |
mistralai/mistral-large-2512Mistral: Mistral Large 3 2512 Mistral's flagship, with a 256K window, vision and tool calling, at a price well under most frontier models. A common European-hosted choice for production reasoning. | VisionTools | 262K | $0.600 | $1.80 |
mistralai/mistral-medium-3Mistral: Mistral Medium 3 The previous Medium generation, at a quarter of Medium 3.5's output price. Often the better value of the two. | VisionTools | 131K | $0.480 | $2.40 |
mistralai/mistral-medium-3-5Mistral: Mistral Medium 3.5 The current Medium tier, priced above Mistral Large on output. Benchmark the two before assuming Medium is the cheaper option. | VisionTools | 262K | $1.80 | $9.00 |
mistralai/mistral-small-2603Mistral: Mistral Small 4 Mistral Small 4, a 256K multimodal tier priced for volume. A sensible default for assistant traffic on Mistral. | VisionTools | 262K | $0.180 | $0.720 |
Anthropic (13)
| Model ID | Capabilities | Context | Input | Output |
|---|---|---|---|---|
anthropic/claude-haiku-4.5Anthropic: Claude Haiku 4.5 The small, fast Claude, built for work where latency and cost matter more than peak reasoning: classification, extraction, routing, and chat that has to answer immediately. | VisionTools | 200K | $1.20 | $6.00 |
anthropic/claude-haiku-4.5Anthropic: Claude Haiku 4.5 The small, fast Claude, built for work where latency and cost matter more than peak reasoning: classification, extraction, routing, and chat that has to answer immediately. | VisionTools | 200K | $1.20 | $6.00 |
anthropic/claude-opus-4.5Anthropic: Claude Opus 4.5 The oldest Opus generation we still serve, retained for reproducibility. Same price as Claude Opus 5, so there is no cost reason to stay on it. | VisionTools | 1M | $6.00 | $30.00 |
anthropic/claude-opus-4.5Anthropic: Claude Opus 4.5 The oldest Opus generation we still serve, retained for reproducibility. Same price as Claude Opus 5, so there is no cost reason to stay on it. | VisionTools | 1M | $6.00 | $30.00 |
anthropic/claude-opus-4.6Anthropic: Claude Opus 4.6 A superseded Opus release, kept so existing pins keep resolving. Anything new should point at Claude Opus 5, which is priced identically. | VisionTools | 1M | $6.00 | $30.00 |
anthropic/claude-opus-4.6Anthropic: Claude Opus 4.6 A superseded Opus release, kept so existing pins keep resolving. Anything new should point at Claude Opus 5, which is priced identically. | VisionTools | 1M | $6.00 | $30.00 |
anthropic/claude-opus-4.7Anthropic: Claude Opus 4.7 An earlier Opus flagship, held in the catalog for pinned integrations. Opus 5 costs the same and is stronger at coding and agent orchestration. | VisionTools | 1M | $6.00 | $30.00 |
anthropic/claude-opus-4.7Anthropic: Claude Opus 4.7 An earlier Opus flagship, held in the catalog for pinned integrations. Opus 5 costs the same and is stronger at coding and agent orchestration. | VisionTools | 1M | $6.00 | $30.00 |
anthropic/claude-opus-4.8Anthropic: Claude Opus 4.8 The last Opus before Claude Opus 5, at the same price. Kept for teams that pinned to it and need reproducible behaviour; new projects should start on Opus 5. | VisionTools | 1M | $6.00 | $30.00 |
anthropic/claude-opus-4.8Anthropic: Claude Opus 4.8 The last Opus before Claude Opus 5, at the same price. Kept for teams that pinned to it and need reproducible behaviour; new projects should start on Opus 5. | VisionTools | 1M | $6.00 | $30.00 |
anthropic/claude-opus-5Anthropic: Claude Opus 5 Anthropic's flagship Opus, built for long-horizon work in real codebases: multi-file features, large refactors, and agent runs that have to stay on track without supervision. Thinking is on by default and it reads dense documents and screenshots at high fidelity. | VisionTools | 1M | $6.00 | $30.00 |
anthropic/claude-sonnet-4.6Anthropic: Claude Sonnet 4.6 The balanced Claude: strong enough for production reasoning and coding, cheap enough to run at volume. The usual default when Opus is more model than the workload needs. | VisionTools | 1M | $3.60 | $18.00 |
anthropic/claude-sonnet-4.6Anthropic: Claude Sonnet 4.6 The balanced Claude: strong enough for production reasoning and coding, cheap enough to run at volume. The usual default when Opus is more model than the workload needs. | VisionTools | 1M | $3.60 | $18.00 |
Qwen (9)
| Model ID | Capabilities | Context | Input | Output |
|---|---|---|---|---|
qwen/qwen3-coder-plusQwen3 Coder Plus A Qwen tuned specifically for code, with a million-token window so it can hold a large repository in context. Aimed at repo-scale refactors and review. | Tools | 1M | $1.20 | $6.00 |
qwen/qwen3-maxQwen3 Max The flagship of the Qwen 3 generation, with a 256K window and tool calling. | Tools | 262K | $0.432 | $1.73 |
qwen/qwen3-vl-flashQwen3 VL Flash The fast vision-language Qwen, and the single cheapest model in this catalog on input tokens. Sized for bulk image and document reading. | VisionTools | 262K | $0.036 | $0.264 |
qwen/qwen3-vl-plusQwen3 VL Plus The vision-language Qwen 3 tier, for reading documents, charts and screenshots over a 256K window. | VisionTools | 262K | $0.180 | $1.73 |
qwen/qwen3.5-flashQwen3.5 Flash One of the cheapest million-token models in the catalog. Built for bulk work over long inputs where per-call cost is the constraint. | Tools | 1M | $0.036 | $0.348 |
qwen/qwen3.5-plusQwen3.5 Plus A million-token Qwen tier priced well below most frontier models. Worth benchmarking when long context matters more than brand. | Tools | 1M | $0.144 | $0.828 |
qwen/qwen3.6-flashQwen3.6 Flash A fast million-token Qwen with vision and tool calling. Positioned for multimodal work that runs constantly. | VisionTools | 1M | $0.300 | $1.80 |
qwen/qwen3.7-maxQwen3.7 Max The top Qwen tier, with a million-token window and tool calling. Alibaba's answer for reasoning-heavy work at frontier scale. | Tools | 1M | $3.00 | $9.00 |
qwen/qwen3.7-plusQwen3.7 Plus The mid Qwen 3.7 tier, adding vision to the million-token window at roughly a fifth of Max's output price. | VisionTools | 1M | $0.480 | $1.92 |
Zai (3)
| Model ID | Capabilities | Context | Input | Output |
|---|---|---|---|---|
zai/glm-4.7Z.AI: GLM 4.7 The previous GLM generation, with a 203K window and tool calling, at roughly two thirds of GLM 5's output price. | Tools | 203K | $0.720 | $2.64 |
zai/glm-4.7-flashZ.AI: GLM 4.7 Flash The fast, cheap GLM tier. One of the least expensive tool-calling models here, sized for high-frequency agent loops. | Tools | 203K | $0.084 | $0.480 |
zai/glm-5Z.AI: GLM 5 Z.AI's current flagship, a 200K tool-calling model positioned for agentic work at open-weight prices. | Tools | 200K | $1.20 | $3.84 |
DeepSeek (2)
| Model ID | Capabilities | Context | Input | Output |
|---|---|---|---|---|
deepseek/deepseek-v4-flashDeepSeek: DeepSeek V4 Flash The fast DeepSeek tier, roughly three times cheaper than Pro while keeping the million-token window. Among the lowest costs per long-context call in the catalog. | Tools | 1.0M | $0.168 | $0.336 |
deepseek/deepseek-v4-proDeepSeek: DeepSeek V4 Pro DeepSeek's flagship, with a million-token window and tool calling at a fraction of Western frontier pricing. A strong value pick for reasoning and code. | Tools | 1.0M | $0.528 | $1.04 |
Moonshot AI (2)
| Model ID | Capabilities | Context | Input | Output |
|---|---|---|---|---|
moonshotai/kimi-k2-thinkingMoonshot AI: Kimi K2 Thinking The reasoning variant of Kimi K2, which works through a problem before answering. Cheaper on output than K2.5 and aimed at analysis rather than chat. | Tools | 256K | $0.720 | $3.00 |
moonshotai/kimi-k2.5Moonshot AI: Kimi K2.5 Moonshot's flagship, a 256K tool-calling model that has become a common open-weight choice for agent work. | Tools | 256K | $0.720 | $3.60 |
MiniMax (1)
| Model ID | Capabilities | Context | Input | Output |
|---|---|---|---|---|
minimax/minimax-m2.5MiniMax M2.5 MiniMax's text flagship, a 196K tool-calling model priced for volume. | Tools | 196K | $0.360 | $1.44 |
This catalog updates automatically. Retrieve it programmatically via the List Models API endpoint.