ApexApiApexApi
Sign In

Supported Models

179 models from 21 model makers. Call any model by its ID in the maker/model format.

Token prices are per 1M tokens; image/video prices are per image/second. All include our service fee.

92 models

OpenAI (33)

Model IDInput
openai/gpt-3.5-turbo

OpenAI: GPT-3.5 Turbo

The model that made cheap chat ordinary. Long since beaten on price and quality by small GPT-5 tiers; kept for legacy integrations.

$0.600
openai/gpt-3.5-turbo-16k

OpenAI: GPT-3.5 Turbo 16k

The extended-context GPT-3.5 variant, now priced above its own successors. Legacy only.

$3.60
openai/gpt-4

OpenAI: GPT-4

The original GPT-4, with an 8K window and the highest token price in the catalog. Present for legacy compatibility, not for new work.

$36.00
openai/gpt-4-turbo

OpenAI: GPT-4 Turbo

The 128K GPT-4 generation that preceded 4o. Superseded on both price and speed, kept for older integrations.

$12.00
openai/gpt-4.1

OpenAI: GPT-4.1

A million-token GPT-4 generation with vision and tool calling, and a long-standing default for document-heavy work. Later GPT-5 tiers usually beat it on price for the same job.

$2.40
openai/gpt-4.1-mini

OpenAI: GPT-4.1 Mini

The mid-size GPT-4.1, keeping the million-token window at a fifth of the output price. A practical choice for long-context work on a budget.

$0.480
openai/gpt-4.1-nano

OpenAI: GPT-4.1 Nano

The smallest GPT-4.1, and one of the cheapest models here that still reads a million tokens. Suited to bulk summarisation and extraction over large corpora.

$0.120
openai/gpt-4o

OpenAI: GPT-4o

OpenAI's multimodal workhorse from the 4-series: text, vision and tool calling over a 128K window. Widely integrated and well understood, though GPT-5 tiers are cheaper for most new work.

$3.00
openai/gpt-4o-2024-05-13

OpenAI: GPT-4o (2024-05-13)

The first GPT-4o snapshot, priced above later ones. Only worth calling if you are reproducing results from that exact version.

$6.00
openai/gpt-4o-2024-08-06

OpenAI: GPT-4o (2024-08-06)

A pinned GPT-4o snapshot, the one that introduced structured outputs. Kept for integrations that were validated against it.

$3.00
openai/gpt-4o-2024-11-20

OpenAI: GPT-4o (2024-11-20)

A pinned GPT-4o snapshot. Use it when you need output that cannot shift as OpenAI updates the rolling alias; otherwise call gpt-4o.

$3.00
openai/gpt-4o-mini

OpenAI: GPT-4o-mini

The small GPT-4o, priced for volume while keeping vision and tool calling. A common default for assistants that must stay cheap per turn.

$0.180
openai/gpt-4o-mini-2024-07-18

OpenAI: GPT-4o-mini (2024-07-18)

A pinned GPT-4o-mini snapshot, for reproducible output from the small 4o tier.

$0.180
openai/gpt-4o-mini-search-preview

OpenAI: GPT-4o-mini Search Preview

The small search-enabled GPT-4o, for grounding high-volume answers in current web results without paying full 4o prices.

$0.180
openai/gpt-4o-search-preview

OpenAI: GPT-4o Search Preview

GPT-4o wired to OpenAI's web search, for answers that need current information rather than training-cutoff knowledge. No tool calling: search is the tool.

$3.00
openai/gpt-5

OpenAI: GPT-5

The first GPT-5 generation, with a 400K context window, vision and tool calling. Still a solid general-purpose default, though later 5.x tiers are usually the better buy.

$1.50
openai/gpt-5-mini

OpenAI: GPT-5 Mini

The mid-size GPT-5, five times cheaper than the full model on output. The usual pick for chat and tool loops that run at volume.

$0.300
openai/gpt-5-nano

OpenAI: GPT-5 Nano

The cheapest GPT-5 tier by a wide margin. Built for classification, routing and short structured extraction at scale.

$0.060
openai/gpt-5.1

OpenAI: GPT-5.1

An incremental GPT-5 release at the same price as GPT-5, with a 400K window. Useful as a pinned target when you need output that does not shift under you.

$1.50
openai/gpt-5.2

OpenAI: GPT-5.2

A GPT-5 generation with a 400K context window, positioned between 5.1 and 5.4. Kept for integrations pinned to it.

$2.10
openai/gpt-5.4

OpenAI: GPT-5.4

A million-token GPT-5 generation at half the output price of 5.5. Worth benchmarking against 5.5 before paying the difference.

$3.00
openai/gpt-5.4-mini

OpenAI: GPT-5.4 Mini

The mid-size GPT-5.4, six times cheaper than the full model on output. Sized for assistant traffic and tool-calling loops that run constantly.

$0.900
openai/gpt-5.4-nano

OpenAI: GPT-5.4 Nano

The smallest GPT-5.4, priced for very high volume: routing, tagging, short extraction, and anything you call thousands of times an hour.

$0.240
openai/gpt-5.5

OpenAI: GPT-5.5

OpenAI's broad multimodal frontier model with a million-token context. A generalist pick when a workload mixes long documents, images and tool use rather than specialising in one.

$6.00
openai/gpt-5.6-luna

OpenAI: GPT-5.6 Luna

The economical GPT-5.6 tier, five times cheaper than Sol on output while keeping the million-token window. Sized for high-volume calls that still need frontier-family behaviour.

$1.20
openai/gpt-5.6-sol

OpenAI: GPT-5.6 Sol

The top tier of the GPT-5.6 family, priced for work where answer quality decides the outcome rather than throughput. Million-token context with vision and tool calling.

$6.00
openai/gpt-5.6-terra

OpenAI: GPT-5.6 Terra

The middle GPT-5.6 tier, at half the price of Sol and twice that of Luna. The tier to try first when you want 5.6 behaviour on production traffic.

$3.00
openai/gpt-oss-120b

OpenAI: GPT-OSS 120B

OpenAI's large open-weight model, served here so you can call it without hosting it. Tool calling included; no vision.

$0.180
openai/gpt-oss-20b

OpenAI: GPT-OSS 20B

The small open-weight OpenAI model and one of the cheapest tool-calling options in the catalog. Good fit for agent loops where per-step cost dominates.

$0.084
openai/o1

OpenAI: o1

OpenAI's first reasoning model, which thinks before answering rather than streaming an immediate reply. Later o-series releases deliver the same approach far more cheaply.

$18.00
openai/o3

OpenAI: o3

A reasoning model for problems where a considered answer beats a fast one: maths, analysis, multi-step planning. Roughly a seventh of o1's output price.

$2.40
openai/o3-mini

OpenAI: o3 Mini

The small o3, for reasoning-shaped work that has to run at volume. Cheaper per call than most frontier chat models.

$1.32
openai/o4-mini

OpenAI: o4 Mini

A compact reasoning model at the same price as o3-mini, worth A/B testing against it on your own tasks before committing.

$1.32

Google (15)

Model IDInput
google/gemini-2.5-flash

Google: Gemini 2.5 Flash

The Gemini 2.5 workhorse: a million tokens of context at a tenth of Pro's price. Still one of the better value picks for long-context summarisation.

$0.360
google/gemini-2.5-flash-image

Google: Nano Banana (Gemini 2.5 Flash Image)

Gemini 2.5 Flash with image understanding, known as Nano Banana. A chat model that reads images rather than an image generator, on a 32K window.

$0.360
google/gemini-2.5-flash-lite

Google: Gemini 2.5 Flash Lite

The cheapest Gemini 2.5 tier, and among the lowest token prices in the whole catalog. Sized for very high-volume, low-complexity calls.

$0.120
google/gemini-2.5-pro

Google: Gemini 2.5 Pro

The Gemini 2.5 flagship, with a million-token window, vision and tool calling. A strong long-document model, now undercut on price by the Gemini 3 line.

$1.50
google/gemini-3-flash-preview

Google: Gemini 3 Flash Preview

The Gemini 3 Flash preview, a million-token multimodal tier priced between 2.5 Flash and 3.5 Flash.

$0.600
google/gemini-3.1-flash-lite

Google: Gemini 3.1 Flash Lite

The cheapest Gemini 3.1 tier that still reads a million tokens. Built for bulk classification and summarisation where per-call cost decides the architecture.

$0.300
google/gemini-3.1-flash-lite-preview

Google: Gemini 3.1 Flash Lite Preview

The preview channel for Gemini 3.1 Flash Lite. Same shape as the stable tier; use it to test upcoming behaviour, not to serve production traffic.

$0.300
google/gemini-3.1-pro-preview

Google: Gemini 3.1 Pro Preview

The Pro tier of Gemini 3.1, for reasoning-heavy work over very long inputs. Preview, so expect behaviour to move before general availability.

$2.40
google/gemini-3.1-pro-preview-customtools

Google: Gemini 3.1 Pro Preview Custom Tools

Gemini 3.1 Pro with custom tool definitions enabled, for agents that call your own functions rather than Google's built-ins.

$2.40
google/gemini-3.5-flash

Google: Gemini 3.5 Flash

The current Flash generation: a million-token window, vision and tool calling at a fraction of Pro pricing. Google's usual answer for high-volume multimodal work.

$1.80
google/gemini-3.5-flash-lite

Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite is Flash-Lite's first step into the agentic space, prioritizing thinking and tool calling to serve as a quick, efficient, and capable subagent for larger, more complex workflows. Compared to Gemini 3.1 Flash-Lite, it delivers stronger agentic performance and tool use, more precise and complete document understanding,

$0.345
google/gemini-3.6-flash

Google: Gemini 3.6 Flash

Gemini 3.6 Flash is designed to deliver strong coding and general-agentic capabilities (near-Pro level) at substantial speed and value with improved token-efficiency and quality over previous Flash models. Gemini 3.6 Flash is optimized for multi-step orchestration, full-stack code refactoring, and general reasoning with much better token efficiency than Gemini 3.5 Flash.

$1.73
google/gemini-3.7-flash

Google: Gemini 3.7 Flash

Gemini 3.7 Flash is Google's newest Flash generation model, aimed at coding and agentic workflows, with a 1M token context window and up to 64k output tokens. On ApexApi it runs as a text in, text out model at the same published price as Gemini 3.6 Flash, so it is the default pick when you want the newer model at that price point. Choose google/gemini-3.6-flash instead when your request needs image input or function calling.

$1.73
google/gemma-4-26b-a4b-it

Google: Gemma 4 26B A4B

The sparse 26B Gemma 4, which activates a fraction of its parameters per token and prices accordingly. One of the cheapest vision-capable models here.

$0.072
google/gemma-4-31b-it

Google: Gemma 4 31B

Google's open-weight instruction-tuned model at 31B, served so you do not have to host it. Multimodal input and tool calling at open-weight pricing.

$0.144

Mistral (14)

Model IDInput
mistralai/codestral-2508

Mistral: Codestral 2508

Mistral's code model, tuned for completion, refactoring and review rather than open-ended chat. 256K of context for repository-scale work.

$0.360
mistralai/devstral-2512

Mistral: Devstral 2 2512

Mistral's agentic coding model, aimed at multi-step development tasks rather than single completions.

$0.480
mistralai/devstral-2512

Mistral: Devstral 2 2512

Mistral's agentic coding model, aimed at multi-step development tasks rather than single completions.

$0.480
mistralai/ministral-14b-2512

Mistral: Ministral 3 14B 2512

The largest Ministral, with symmetric input and output pricing that makes cost easy to predict for generation-heavy work.

$0.240
mistralai/ministral-14b-2512

Mistral: Ministral 3 14B 2512

The largest Ministral, with symmetric input and output pricing that makes cost easy to predict for generation-heavy work.

$0.240
mistralai/ministral-3b-2512

Mistral: Ministral 3 3B 2512

The smallest Ministral, and one of the cheapest models here. Built for edge-shaped workloads: routing, tagging, short replies.

$0.120
mistralai/ministral-3b-2512

Mistral: Ministral 3 3B 2512

The smallest Ministral, and one of the cheapest models here. Built for edge-shaped workloads: routing, tagging, short replies.

$0.120
mistralai/ministral-8b-2512

Mistral: Ministral 3 8B 2512

The mid Ministral at 8B, multimodal with tool calling, priced flat on input and output.

$0.180
mistralai/ministral-8b-2512

Mistral: Ministral 3 8B 2512

The mid Ministral at 8B, multimodal with tool calling, priced flat on input and output.

$0.180
mistralai/mistral-large-2512

Mistral: Mistral Large 3 2512

Mistral's flagship, with a 256K window, vision and tool calling, at a price well under most frontier models. A common European-hosted choice for production reasoning.

$0.600
mistralai/mistral-large-2512

Mistral: Mistral Large 3 2512

Mistral's flagship, with a 256K window, vision and tool calling, at a price well under most frontier models. A common European-hosted choice for production reasoning.

$0.600
mistralai/mistral-medium-3

Mistral: Mistral Medium 3

The previous Medium generation, at a quarter of Medium 3.5's output price. Often the better value of the two.

$0.480
mistralai/mistral-medium-3-5

Mistral: Mistral Medium 3.5

The current Medium tier, priced above Mistral Large on output. Benchmark the two before assuming Medium is the cheaper option.

$1.80
mistralai/mistral-small-2603

Mistral: Mistral Small 4

Mistral Small 4, a 256K multimodal tier priced for volume. A sensible default for assistant traffic on Mistral.

$0.180

Anthropic (13)

Model IDInput
anthropic/claude-haiku-4.5

Anthropic: Claude Haiku 4.5

The small, fast Claude, built for work where latency and cost matter more than peak reasoning: classification, extraction, routing, and chat that has to answer immediately.

$1.20
anthropic/claude-haiku-4.5

Anthropic: Claude Haiku 4.5

The small, fast Claude, built for work where latency and cost matter more than peak reasoning: classification, extraction, routing, and chat that has to answer immediately.

$1.20
anthropic/claude-opus-4.5

Anthropic: Claude Opus 4.5

The oldest Opus generation we still serve, retained for reproducibility. Same price as Claude Opus 5, so there is no cost reason to stay on it.

$6.00
anthropic/claude-opus-4.5

Anthropic: Claude Opus 4.5

The oldest Opus generation we still serve, retained for reproducibility. Same price as Claude Opus 5, so there is no cost reason to stay on it.

$6.00
anthropic/claude-opus-4.6

Anthropic: Claude Opus 4.6

A superseded Opus release, kept so existing pins keep resolving. Anything new should point at Claude Opus 5, which is priced identically.

$6.00
anthropic/claude-opus-4.6

Anthropic: Claude Opus 4.6

A superseded Opus release, kept so existing pins keep resolving. Anything new should point at Claude Opus 5, which is priced identically.

$6.00
anthropic/claude-opus-4.7

Anthropic: Claude Opus 4.7

An earlier Opus flagship, held in the catalog for pinned integrations. Opus 5 costs the same and is stronger at coding and agent orchestration.

$6.00
anthropic/claude-opus-4.7

Anthropic: Claude Opus 4.7

An earlier Opus flagship, held in the catalog for pinned integrations. Opus 5 costs the same and is stronger at coding and agent orchestration.

$6.00
anthropic/claude-opus-4.8

Anthropic: Claude Opus 4.8

The last Opus before Claude Opus 5, at the same price. Kept for teams that pinned to it and need reproducible behaviour; new projects should start on Opus 5.

$6.00
anthropic/claude-opus-4.8

Anthropic: Claude Opus 4.8

The last Opus before Claude Opus 5, at the same price. Kept for teams that pinned to it and need reproducible behaviour; new projects should start on Opus 5.

$6.00
anthropic/claude-opus-5

Anthropic: Claude Opus 5

Anthropic's flagship Opus, built for long-horizon work in real codebases: multi-file features, large refactors, and agent runs that have to stay on track without supervision. Thinking is on by default and it reads dense documents and screenshots at high fidelity.

$6.00
anthropic/claude-sonnet-4.6

Anthropic: Claude Sonnet 4.6

The balanced Claude: strong enough for production reasoning and coding, cheap enough to run at volume. The usual default when Opus is more model than the workload needs.

$3.60
anthropic/claude-sonnet-4.6

Anthropic: Claude Sonnet 4.6

The balanced Claude: strong enough for production reasoning and coding, cheap enough to run at volume. The usual default when Opus is more model than the workload needs.

$3.60

Qwen (9)

Model IDInput
qwen/qwen3-coder-plus

Qwen3 Coder Plus

A Qwen tuned specifically for code, with a million-token window so it can hold a large repository in context. Aimed at repo-scale refactors and review.

$1.20
qwen/qwen3-max

Qwen3 Max

The flagship of the Qwen 3 generation, with a 256K window and tool calling.

$0.432
qwen/qwen3-vl-flash

Qwen3 VL Flash

The fast vision-language Qwen, and the single cheapest model in this catalog on input tokens. Sized for bulk image and document reading.

$0.036
qwen/qwen3-vl-plus

Qwen3 VL Plus

The vision-language Qwen 3 tier, for reading documents, charts and screenshots over a 256K window.

$0.180
qwen/qwen3.5-flash

Qwen3.5 Flash

One of the cheapest million-token models in the catalog. Built for bulk work over long inputs where per-call cost is the constraint.

$0.036
qwen/qwen3.5-plus

Qwen3.5 Plus

A million-token Qwen tier priced well below most frontier models. Worth benchmarking when long context matters more than brand.

$0.144
qwen/qwen3.6-flash

Qwen3.6 Flash

A fast million-token Qwen with vision and tool calling. Positioned for multimodal work that runs constantly.

$0.300
qwen/qwen3.7-max

Qwen3.7 Max

The top Qwen tier, with a million-token window and tool calling. Alibaba's answer for reasoning-heavy work at frontier scale.

$3.00
qwen/qwen3.7-plus

Qwen3.7 Plus

The mid Qwen 3.7 tier, adding vision to the million-token window at roughly a fifth of Max's output price.

$0.480

Zai (3)

Model IDInput
zai/glm-4.7

Z.AI: GLM 4.7

The previous GLM generation, with a 203K window and tool calling, at roughly two thirds of GLM 5's output price.

$0.720
zai/glm-4.7-flash

Z.AI: GLM 4.7 Flash

The fast, cheap GLM tier. One of the least expensive tool-calling models here, sized for high-frequency agent loops.

$0.084
zai/glm-5

Z.AI: GLM 5

Z.AI's current flagship, a 200K tool-calling model positioned for agentic work at open-weight prices.

$1.20

DeepSeek (2)

Model IDInput
deepseek/deepseek-v4-flash

DeepSeek: DeepSeek V4 Flash

The fast DeepSeek tier, roughly three times cheaper than Pro while keeping the million-token window. Among the lowest costs per long-context call in the catalog.

$0.168
deepseek/deepseek-v4-pro

DeepSeek: DeepSeek V4 Pro

DeepSeek's flagship, with a million-token window and tool calling at a fraction of Western frontier pricing. A strong value pick for reasoning and code.

$0.528

Moonshot AI (2)

Model IDInput
moonshotai/kimi-k2-thinking

Moonshot AI: Kimi K2 Thinking

The reasoning variant of Kimi K2, which works through a problem before answering. Cheaper on output than K2.5 and aimed at analysis rather than chat.

$0.720
moonshotai/kimi-k2.5

Moonshot AI: Kimi K2.5

Moonshot's flagship, a 256K tool-calling model that has become a common open-weight choice for agent work.

$0.720

MiniMax (1)

Model IDInput
minimax/minimax-m2.5

MiniMax M2.5

MiniMax's text flagship, a 196K tool-calling model priced for volume.

$0.360

This catalog updates automatically. Retrieve it programmatically via the List Models API endpoint.