ApexApiApexApi
Sign In

AI Model Catalog

145 models from 21 model makers, all through a single API.

Token prices are per 1M tokens; image/video prices are per image/second. All include our service fee.

OpenAI

34 models
ModelInput
openai/gpt-3.5-turbo

OpenAI: GPT-3.5 Turbo

The model that made cheap chat ordinary. Long since beaten on price and quality by small GPT-5 tiers; kept for legacy integrations.

$0.600
openai/gpt-3.5-turbo-16k

OpenAI: GPT-3.5 Turbo 16k

The extended-context GPT-3.5 variant, now priced above its own successors. Legacy only.

$3.60
openai/gpt-4

OpenAI: GPT-4

The original GPT-4, with an 8K window and the highest token price in the catalog. Present for legacy compatibility, not for new work.

$36.00
openai/gpt-4-turbo

OpenAI: GPT-4 Turbo

The 128K GPT-4 generation that preceded 4o. Superseded on both price and speed, kept for older integrations.

$12.00
openai/gpt-4.1

OpenAI: GPT-4.1

A million-token GPT-4 generation with vision and tool calling, and a long-standing default for document-heavy work. Later GPT-5 tiers usually beat it on price for the same job.

$2.40
openai/gpt-4.1-mini

OpenAI: GPT-4.1 Mini

The mid-size GPT-4.1, keeping the million-token window at a fifth of the output price. A practical choice for long-context work on a budget.

$0.480
openai/gpt-4.1-nano

OpenAI: GPT-4.1 Nano

The smallest GPT-4.1, and one of the cheapest models here that still reads a million tokens. Suited to bulk summarisation and extraction over large corpora.

$0.120
openai/gpt-4o

OpenAI: GPT-4o

OpenAI's multimodal workhorse from the 4-series: text, vision and tool calling over a 128K window. Widely integrated and well understood, though GPT-5 tiers are cheaper for most new work.

$3.00
openai/gpt-4o-2024-05-13

OpenAI: GPT-4o (2024-05-13)

The first GPT-4o snapshot, priced above later ones. Only worth calling if you are reproducing results from that exact version.

$6.00
openai/gpt-4o-2024-08-06

OpenAI: GPT-4o (2024-08-06)

A pinned GPT-4o snapshot, the one that introduced structured outputs. Kept for integrations that were validated against it.

$3.00
openai/gpt-4o-2024-11-20

OpenAI: GPT-4o (2024-11-20)

A pinned GPT-4o snapshot. Use it when you need output that cannot shift as OpenAI updates the rolling alias; otherwise call gpt-4o.

$3.00
openai/gpt-4o-mini

OpenAI: GPT-4o-mini

The small GPT-4o, priced for volume while keeping vision and tool calling. A common default for assistants that must stay cheap per turn.

$0.180
openai/gpt-4o-mini-2024-07-18

OpenAI: GPT-4o-mini (2024-07-18)

A pinned GPT-4o-mini snapshot, for reproducible output from the small 4o tier.

$0.180
openai/gpt-4o-mini-search-preview

OpenAI: GPT-4o-mini Search Preview

The small search-enabled GPT-4o, for grounding high-volume answers in current web results without paying full 4o prices.

$0.180
openai/gpt-4o-search-preview

OpenAI: GPT-4o Search Preview

GPT-4o wired to OpenAI's web search, for answers that need current information rather than training-cutoff knowledge. No tool calling: search is the tool.

$3.00
openai/gpt-5

OpenAI: GPT-5

The first GPT-5 generation, with a 400K context window, vision and tool calling. Still a solid general-purpose default, though later 5.x tiers are usually the better buy.

$1.50
openai/gpt-5-mini

OpenAI: GPT-5 Mini

The mid-size GPT-5, five times cheaper than the full model on output. The usual pick for chat and tool loops that run at volume.

$0.300
openai/gpt-5-nano

OpenAI: GPT-5 Nano

The cheapest GPT-5 tier by a wide margin. Built for classification, routing and short structured extraction at scale.

$0.060
openai/gpt-5.1

OpenAI: GPT-5.1

An incremental GPT-5 release at the same price as GPT-5, with a 400K window. Useful as a pinned target when you need output that does not shift under you.

$1.50
openai/gpt-5.2

OpenAI: GPT-5.2

A GPT-5 generation with a 400K context window, positioned between 5.1 and 5.4. Kept for integrations pinned to it.

$2.10
openai/gpt-5.4

OpenAI: GPT-5.4

A million-token GPT-5 generation at half the output price of 5.5. Worth benchmarking against 5.5 before paying the difference.

$3.00
openai/gpt-5.4-mini

OpenAI: GPT-5.4 Mini

The mid-size GPT-5.4, six times cheaper than the full model on output. Sized for assistant traffic and tool-calling loops that run constantly.

$0.900
openai/gpt-5.4-nano

OpenAI: GPT-5.4 Nano

The smallest GPT-5.4, priced for very high volume: routing, tagging, short extraction, and anything you call thousands of times an hour.

$0.240
openai/gpt-5.5

OpenAI: GPT-5.5

OpenAI's broad multimodal frontier model with a million-token context. A generalist pick when a workload mixes long documents, images and tool use rather than specialising in one.

$6.00
openai/gpt-5.6-luna

OpenAI: GPT-5.6 Luna

The economical GPT-5.6 tier, five times cheaper than Sol on output while keeping the million-token window. Sized for high-volume calls that still need frontier-family behaviour.

$1.20
openai/gpt-5.6-sol

OpenAI: GPT-5.6 Sol

The top tier of the GPT-5.6 family, priced for work where answer quality decides the outcome rather than throughput. Million-token context with vision and tool calling.

$6.00
openai/gpt-5.6-terra

OpenAI: GPT-5.6 Terra

The middle GPT-5.6 tier, at half the price of Sol and twice that of Luna. The tier to try first when you want 5.6 behaviour on production traffic.

$3.00
openai/gpt-image-2

GPT Image 2

OpenAI's image model, notable for following long, detailed prompts and rendering readable text inside the image.

openai/gpt-oss-120b

OpenAI: GPT-OSS 120B

OpenAI's large open-weight model, served here so you can call it without hosting it. Tool calling included; no vision.

$0.180
openai/gpt-oss-20b

OpenAI: GPT-OSS 20B

The small open-weight OpenAI model and one of the cheapest tool-calling options in the catalog. Good fit for agent loops where per-step cost dominates.

$0.084
openai/o1

OpenAI: o1

OpenAI's first reasoning model, which thinks before answering rather than streaming an immediate reply. Later o-series releases deliver the same approach far more cheaply.

$18.00
openai/o3

OpenAI: o3

A reasoning model for problems where a considered answer beats a fast one: maths, analysis, multi-step planning. Roughly a seventh of o1's output price.

$2.40
openai/o3-mini

OpenAI: o3 Mini

The small o3, for reasoning-shaped work that has to run at volume. Cheaper per call than most frontier chat models.

$1.32
openai/o4-mini

OpenAI: o4 Mini

A compact reasoning model at the same price as o3-mini, worth A/B testing against it on your own tasks before committing.

$1.32

Google

33 models
ModelInput
google/gemini-2.5-flash

Google: Gemini 2.5 Flash

The Gemini 2.5 workhorse: a million tokens of context at a tenth of Pro's price. Still one of the better value picks for long-context summarisation.

$0.360
google/gemini-2.5-flash-image

Google: Nano Banana (Gemini 2.5 Flash Image)

Gemini 2.5 Flash with image understanding, known as Nano Banana. A chat model that reads images rather than an image generator, on a 32K window.

$0.360
google/gemini-2.5-flash-lite

Google: Gemini 2.5 Flash Lite

The cheapest Gemini 2.5 tier, and among the lowest token prices in the whole catalog. Sized for very high-volume, low-complexity calls.

$0.120
google/gemini-2.5-pro

Google: Gemini 2.5 Pro

The Gemini 2.5 flagship, with a million-token window, vision and tool calling. A strong long-document model, now undercut on price by the Gemini 3 line.

$1.50
google/gemini-3-flash-preview

Google: Gemini 3 Flash Preview

The Gemini 3 Flash preview, a million-token multimodal tier priced between 2.5 Flash and 3.5 Flash.

$0.600
google/gemini-3-pro-image-preview

Gemini 3 Pro Image Preview (Vertex)

Gemini 3 Pro's image generator, aimed at complex prompts and legible in-image text. Preview channel, so output may shift.

google/gemini-3-pro-image-preview/edit

Gemini 3 Pro Image — Edit

Gemini 3 Pro Image in edit mode: pass a source image and an instruction rather than a prompt alone.

google/gemini-3.1-flash-image-preview

Gemini 3.1 Flash Image (Preview) — direct

The fast Gemini 3.1 image tier, roughly half the price of Gemini 3 Pro Image. Suited to volume generation where turnaround beats maximum fidelity.

google/gemini-3.1-flash-image-preview/edit

Gemini 3.1 Flash Image — Edit

Gemini 3.1 Flash Image in edit mode, for quick instruction-based revisions of an existing image.

google/gemini-3.1-flash-lite

Google: Gemini 3.1 Flash Lite

The cheapest Gemini 3.1 tier that still reads a million tokens. Built for bulk classification and summarisation where per-call cost decides the architecture.

$0.300
google/gemini-3.1-flash-lite-preview

Google: Gemini 3.1 Flash Lite Preview

The preview channel for Gemini 3.1 Flash Lite. Same shape as the stable tier; use it to test upcoming behaviour, not to serve production traffic.

$0.300
google/gemini-3.1-pro-preview

Google: Gemini 3.1 Pro Preview

The Pro tier of Gemini 3.1, for reasoning-heavy work over very long inputs. Preview, so expect behaviour to move before general availability.

$2.40
google/gemini-3.1-pro-preview-customtools

Google: Gemini 3.1 Pro Preview Custom Tools

Gemini 3.1 Pro with custom tool definitions enabled, for agents that call your own functions rather than Google's built-ins.

$2.40
google/gemini-3.5-flash

Google: Gemini 3.5 Flash

The current Flash generation: a million-token window, vision and tool calling at a fraction of Pro pricing. Google's usual answer for high-volume multimodal work.

$1.80
google/gemini-3.5-flash-lite

Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite is Flash-Lite's first step into the agentic space, prioritizing thinking and tool calling to serve as a quick, efficient, and capable subagent for larger, more complex workflows. Compared to Gemini 3.1 Flash-Lite, it delivers stronger agentic performance and tool use, more precise and complete document understanding,

$0.345
google/gemini-3.6-flash

Google: Gemini 3.6 Flash

Gemini 3.6 Flash is designed to deliver strong coding and general-agentic capabilities (near-Pro level) at substantial speed and value with improved token-efficiency and quality over previous Flash models. Gemini 3.6 Flash is optimized for multi-step orchestration, full-stack code refactoring, and general reasoning with much better token efficiency than Gemini 3.5 Flash.

$1.73
google/gemini-3.7-flash

Google: Gemini 3.7 Flash

Gemini 3.7 Flash is Google's newest Flash generation model, aimed at coding and agentic workflows, with a 1M token context window and up to 64k output tokens. On ApexApi it runs as a text in, text out model at the same published price as Gemini 3.6 Flash, so it is the default pick when you want the newer model at that price point. Choose google/gemini-3.6-flash instead when your request needs image input or function calling.

$1.73
google/gemini-flash-image

Gemini 2.5 Flash Image (Nano Banana)

The original Nano Banana, Gemini 2.5 Flash Image. Cheapest of the Google image tiers and still capable for straightforward generation.

google/gemini-omni-flash/edit

Gemini Omni Flash - Video Edit

Edit an existing video with a natural language instruction. Provide a source video up to 10 seconds and describe the change, for example "Make this video anime. Keep everything else the same." The output keeps the scene and length of the source. 720p output with audio.

google/gemini-omni-flash/image-to-video

Gemini Omni Flash - Image to Video

Gemini Omni Flash image-to-video animates a still image into a 3-10 second 720p clip at 24 fps, guided by your text prompt, with generated audio included. Supply a starting frame and describe the motion; the model preserves the look of the source image while adding coherent movement and sound.

google/gemini-omni-flash/text-to-video

Gemini Omni Flash - Text to Video

Gemini Omni Flash is Google's conversational video generation model. It turns a text prompt into a 3-10 second 720p clip at 24 fps with a generated audio track included on every render. Built on the same Omni model that powers video editing in Gemini, it is tuned for fast turnaround and cinematic camera direction from plain-language prompts.

google/gemma-4-26b-a4b-it

Google: Gemma 4 26B A4B

The sparse 26B Gemma 4, which activates a fraction of its parameters per token and prices accordingly. One of the cheapest vision-capable models here.

$0.072
google/gemma-4-31b-it

Google: Gemma 4 31B

Google's open-weight instruction-tuned model at 31B, served so you do not have to host it. Multimodal input and tool calling at open-weight pricing.

$0.144
google/nano-banana-2

Nano Banana 2

The mid Nano Banana tier at roughly half Pro's price. A good default for drafts and iteration before committing to a Pro render.

google/nano-banana-2/edit

Nano Banana 2 Edit

Nano Banana 2 in edit mode, for instruction-based changes to an existing image at half the cost of the Pro editor.

google/nano-banana-pro

Nano Banana Pro

Google's top image model, the one to reach for when prompt adherence and text rendering inside the image actually matter. The priciest image tier here, and usually worth it for final assets.

google/nano-banana-pro/edit

Nano Banana Pro Edit

Nano Banana Pro in edit mode: supply an image and describe the change. Best of the Google editors at keeping the untouched parts of a picture intact.

google/veo3

Google Veo 3

The previous Veo generation, and the most expensive video model in the catalog. Veo 3.1 costs less than half as much; only stay here for output you have already validated.

google/veo3.1

Veo 3.1

Google's current Veo generation, generating video with synchronised audio from a text prompt. The default choice here when the clip has to look finished rather than rough.

google/veo3.1/fast

Veo 3.1 Fast

Veo 3.1 at half price, trading some fidelity for turnaround. The tier to iterate on before committing to a full-price render.

google/veo3.1/fast/image-to-video

Veo 3.1 Fast — Image to Video

The fast Veo 3.1 tier driven by a source image. Cheapest way to test how a still will animate before paying for the full model.

google/veo3.1/image-to-video

Veo 3.1 — Image to Video

Veo 3.1 driven by a still image: supply a picture as the first frame and describe how it should move. Use this when the look is already locked and only the motion is in question.

google/veo3.1/lite

Veo 3.1 Lite

The cheapest Veo tier, a quarter the price of standard 3.1. Built for volume drafts and previews.

Bytedance

14 models
ModelInput
bytedance/seedance-2.0/fast/image-to-video

Seedance 2.0 Fast — Image to Video

The fast Seedance 2.0 tier animating a supplied first frame. Cheaper turnaround while you settle on the motion.

bytedance/seedance-2.0/fast/reference-to-video

Seedance 2.0 Fast — Reference to Video

The fast Seedance 2.0 tier working from reference images, for testing subject consistency before a full-price render.

bytedance/seedance-2.0/fast/text-to-video

Seedance 2.0 Fast — Text to Video

The fast Seedance 2.0 tier generating from a text prompt, at four fifths of standard price. Sized for iterating on ideas.

bytedance/seedance-2.0/image-to-video

Seedance 2.0 — Image to Video

Seedance 2.0 animating a still you supply as the first frame. Use it when the composition is settled and you only need motion.

bytedance/seedance-2.0/mini/image-to-video

Seedance 2.0 Mini — Image to Video

The mini Seedance tier animating a supplied still, for quick motion tests at the lowest price in the family.

bytedance/seedance-2.0/mini/reference-to-video

Seedance 2.0 Mini — Reference to Video

The mini Seedance tier driven by reference images, for cheap checks on whether a subject holds up across a shot.

bytedance/seedance-2.0/mini/text-to-video

Seedance 2.0 Mini — Text to Video

The cheapest Seedance tier generating from text, at half the price of standard 2.0. Built for volume drafts.

bytedance/seedance-2.0/reference-to-video

Seedance 2.0 — Reference to Video

Seedance 2.0 guided by reference images that fix a character, product or style, rather than by a single first frame. The mode for keeping a subject consistent across shots.

bytedance/seedance-2.0/text-to-video

Seedance 2.0 — Text to Video

Seedance 2.0 generating a clip from a written prompt alone. The starting point when nothing visual exists yet.

bytedance/seedance-2.5/image-to-video

Seedance 2.5 - Image to Video

Animates a still image into up to 30 seconds of video with synchronized audio, keeping the subject consistent across the full clip. Supports an optional end frame to control where the video lands. Choose it over Seedance 2.0 image-to-video for longer takes and native audio.

bytedance/seedance-2.5/reference-to-video

Seedance 2.5 - Reference to Video

Generates video from up to 50 reference files: images for identity and wardrobe, video clips for motion style or editing, audio for rhythm and voice. Reference video seconds are billed together with output seconds at a reduced rate. This is the most controllable Seedance 2.5 variant and the successor to Seedance 2.0 reference-to-video.

bytedance/seedance-2.5/text-to-video

Seedance 2.5 - Text to Video

Seedance 2.5 is ByteDance's newest video model, generating up to 30 seconds of single-shot video with synchronized audio at 480p or 720p. Pick it over Seedance 2.0 when you need clips longer than 15 seconds or native audio that stays in sync with on-screen action. Billing is per token of output video, roughly $0.57 per second at 720p.

bytedance/seedance/v1/pro

Seedance 1.0 Pro

The previous Seedance generation, still less than half the price of Seedance 2.0 standard. Worth keeping for bulk work where 2.0 is more than the shot needs.

bytedance/seedream-5.0

Seedream 5.0 Pro

ByteDance's flagship image model. The most expensive image tier here, aimed at high-fidelity commercial output.

Mistral

9 models
ModelInput
mistralai/codestral-2508

Mistral: Codestral 2508

Mistral's code model, tuned for completion, refactoring and review rather than open-ended chat. 256K of context for repository-scale work.

$0.360
mistralai/devstral-2512

Mistral: Devstral 2 2512

Mistral's agentic coding model, aimed at multi-step development tasks rather than single completions.

$0.480
mistralai/ministral-14b-2512

Mistral: Ministral 3 14B 2512

The largest Ministral, with symmetric input and output pricing that makes cost easy to predict for generation-heavy work.

$0.240
mistralai/ministral-3b-2512

Mistral: Ministral 3 3B 2512

The smallest Ministral, and one of the cheapest models here. Built for edge-shaped workloads: routing, tagging, short replies.

$0.120
mistralai/ministral-8b-2512

Mistral: Ministral 3 8B 2512

The mid Ministral at 8B, multimodal with tool calling, priced flat on input and output.

$0.180
mistralai/mistral-large-2512

Mistral: Mistral Large 3 2512

Mistral's flagship, with a 256K window, vision and tool calling, at a price well under most frontier models. A common European-hosted choice for production reasoning.

$0.600
mistralai/mistral-medium-3

Mistral: Mistral Medium 3

The previous Medium generation, at a quarter of Medium 3.5's output price. Often the better value of the two.

$0.480
mistralai/mistral-medium-3-5

Mistral: Mistral Medium 3.5

The current Medium tier, priced above Mistral Large on output. Benchmark the two before assuming Medium is the cheaper option.

$1.80
mistralai/mistral-small-2603

Mistral: Mistral Small 4

Mistral Small 4, a 256K multimodal tier priced for volume. A sensible default for assistant traffic on Mistral.

$0.180

Qwen

9 models
ModelInput
qwen/qwen3-coder-plus

Qwen3 Coder Plus

A Qwen tuned specifically for code, with a million-token window so it can hold a large repository in context. Aimed at repo-scale refactors and review.

$1.20
qwen/qwen3-max

Qwen3 Max

The flagship of the Qwen 3 generation, with a 256K window and tool calling.

$0.432
qwen/qwen3-vl-flash

Qwen3 VL Flash

The fast vision-language Qwen, and the single cheapest model in this catalog on input tokens. Sized for bulk image and document reading.

$0.036
qwen/qwen3-vl-plus

Qwen3 VL Plus

The vision-language Qwen 3 tier, for reading documents, charts and screenshots over a 256K window.

$0.180
qwen/qwen3.5-flash

Qwen3.5 Flash

One of the cheapest million-token models in the catalog. Built for bulk work over long inputs where per-call cost is the constraint.

$0.036
qwen/qwen3.5-plus

Qwen3.5 Plus

A million-token Qwen tier priced well below most frontier models. Worth benchmarking when long context matters more than brand.

$0.144
qwen/qwen3.6-flash

Qwen3.6 Flash

A fast million-token Qwen with vision and tool calling. Positioned for multimodal work that runs constantly.

$0.300
qwen/qwen3.7-max

Qwen3.7 Max

The top Qwen tier, with a million-token window and tool calling. Alibaba's answer for reasoning-heavy work at frontier scale.

$3.00
qwen/qwen3.7-plus

Qwen3.7 Plus

The mid Qwen 3.7 tier, adding vision to the million-token window at roughly a fifth of Max's output price.

$0.480

Anthropic

7 models
ModelInput
anthropic/claude-haiku-4.5

Anthropic: Claude Haiku 4.5

The small, fast Claude, built for work where latency and cost matter more than peak reasoning: classification, extraction, routing, and chat that has to answer immediately.

$1.20
anthropic/claude-opus-4.5

Anthropic: Claude Opus 4.5

The oldest Opus generation we still serve, retained for reproducibility. Same price as Claude Opus 5, so there is no cost reason to stay on it.

$6.00
anthropic/claude-opus-4.6

Anthropic: Claude Opus 4.6

A superseded Opus release, kept so existing pins keep resolving. Anything new should point at Claude Opus 5, which is priced identically.

$6.00
anthropic/claude-opus-4.7

Anthropic: Claude Opus 4.7

An earlier Opus flagship, held in the catalog for pinned integrations. Opus 5 costs the same and is stronger at coding and agent orchestration.

$6.00
anthropic/claude-opus-4.8

Anthropic: Claude Opus 4.8

The last Opus before Claude Opus 5, at the same price. Kept for teams that pinned to it and need reproducible behaviour; new projects should start on Opus 5.

$6.00
anthropic/claude-opus-5

Anthropic: Claude Opus 5

Anthropic's flagship Opus, built for long-horizon work in real codebases: multi-file features, large refactors, and agent runs that have to stay on track without supervision. Thinking is on by default and it reads dense documents and screenshots at high fidelity.

$6.00
anthropic/claude-sonnet-4.6

Anthropic: Claude Sonnet 4.6

The balanced Claude: strong enough for production reasoning and coding, cheap enough to run at volume. The usual default when Opus is more model than the workload needs.

$3.60

Black Forest Labs

7 models
ModelInput
black-forest-labs/flux-2

FLUX.2

The standard FLUX.2 tier, two and a half times cheaper than Pro. A sensible default for drafts and volume work.

black-forest-labs/flux-2-pro

FLUX.2 Pro

The professional FLUX.2 tier, for final assets where prompt adherence and fine detail matter. The strongest FLUX option in the catalog.

black-forest-labs/flux-2-pro/edit

FLUX.2 Pro Edit

FLUX.2 Pro in edit mode: supply a source image and describe the change, rather than generating from a prompt alone.

black-forest-labs/flux-2/edit

FLUX.2 Edit

FLUX.2 in edit mode, for instruction-based image changes at the standard tier price.

black-forest-labs/flux-pro

FLUX.1 Pro

The FLUX.1 professional tier. Superseded by FLUX.2 Pro, kept for pipelines tuned to its particular look.

black-forest-labs/flux/dev

FLUX.1 Dev

The open-weight FLUX.1 development model, served so you do not have to run it yourself. A common baseline for fine-tuning workflows.

black-forest-labs/flux/schnell

FLUX.1 Schnell

The fastest, cheapest FLUX, distilled for few-step generation. Built for thumbnails, previews and anything you generate by the hundred.

Xai

7 models
ModelInput
xai/grok-imagine-image

Grok Imagine — Image

xAI's image generator at its standard tier, priced low enough for heavy iteration.

xai/grok-imagine-image/edit

Grok Imagine — Image Edit

Grok Imagine in edit mode: pass an image and an instruction to change it, at the standard tier price.

xai/grok-imagine-image/quality/edit

Grok Imagine — Quality Edit

The quality tier of Grok Imagine in edit mode, for revisions where the standard tier loses too much of the original.

xai/grok-imagine-image/quality/text-to-image

Grok Imagine — Quality

The quality tier of Grok Imagine, generating from a text prompt alone. Two and a half times the standard price for a cleaner render.

xai/grok-imagine-video/image-to-video

Grok Imagine — Image to Video

Grok Imagine animating a still you supply as the opening frame.

xai/grok-imagine-video/reference-to-video

Grok Imagine — Reference to Video

Grok Imagine guided by reference images rather than a single first frame, for holding a subject steady across a clip.

xai/grok-imagine-video/text-to-video

Grok Imagine — Text to Video

xAI's video model generating from a text prompt, at one of the lowest per-second prices in the catalog.

Kuaishou

6 models
ModelInput
kuaishou/kling-video/o3/pro/image-to-video

Kling o3 Pro — Image to Video

The o3 Kling line at Pro tier, animating a supplied still. Priced level with Kling 3.0 Pro, so worth A/B testing on your own footage.

kuaishou/kling-video/v2/master/text-to-video

Kling 2.0 Master Text-to-Video

The Kling 2.0 Master tier, generating from text. Two and a half times the price of Kling 3.0 Pro, so only stay here for output you have already signed off.

kuaishou/kling-video/v3/pro/image-to-video

Kling 3.0 Pro — Image to Video

Kling 3.0 Pro animating a still you supply. The mode to use when the frame is already art-directed.

kuaishou/kling-video/v3/pro/motion-control

Kling 3 Pro Motion Control

Kling 3 Pro driven by a reference video: the motion comes from your clip, the subject from your image. Both the image and the driving video need a visible human upper body or the request is rejected.

kuaishou/kling-video/v3/pro/text-to-video

Kling 3.0 Pro — Text to Video

Kling 3.0 Pro generating a clip from a text prompt. Kling is generally the strongest option here for human motion.

kuaishou/kling-video/v3/standard/image-to-video

Kling 3.0 Standard — Image to Video

The standard Kling 3.0 tier animating a supplied image, at three quarters of Pro pricing. A reasonable default before escalating to Pro.

Elevenlabs

3 models
ModelInput
elevenlabs/tts/eleven-v3

ElevenLabs TTS Eleven v3

ElevenLabs' most expressive voice model, for narration and character work where delivery and emotion carry the piece.

elevenlabs/tts/multilingual-v2

ElevenLabs TTS Multilingual v2

The multilingual ElevenLabs voice, for output that has to hold a consistent voice across languages.

elevenlabs/tts/turbo-v2.5

ElevenLabs TTS Turbo v2.5

The fast ElevenLabs tier at half the price of v3. Built for real-time and high-volume speech where latency matters more than performance nuance.

Zai

3 models
ModelInput
zai/glm-4.7

Z.AI: GLM 4.7

The previous GLM generation, with a 203K window and tool calling, at roughly two thirds of GLM 5's output price.

$0.720
zai/glm-4.7-flash

Z.AI: GLM 4.7 Flash

The fast, cheap GLM tier. One of the least expensive tool-calling models here, sized for high-frequency agent loops.

$0.084
zai/glm-5

Z.AI: GLM 5

Z.AI's current flagship, a 200K tool-calling model positioned for agentic work at open-weight prices.

$1.20

DeepSeek

2 models
ModelInput
deepseek/deepseek-v4-flash

DeepSeek: DeepSeek V4 Flash

The fast DeepSeek tier, roughly three times cheaper than Pro while keeping the million-token window. Among the lowest costs per long-context call in the catalog.

$0.168
deepseek/deepseek-v4-pro

DeepSeek: DeepSeek V4 Pro

DeepSeek's flagship, with a million-token window and tool calling at a fraction of Western frontier pricing. A strong value pick for reasoning and code.

$0.528

MiniMax

2 models
ModelInput
minimax/hailuo-02

MiniMax Hailuo-02

MiniMax's video model, priced between the budget and premium tiers here. A solid middle option when Kling is too expensive and the cheapest tiers look it.

minimax/minimax-m2.5

MiniMax M2.5

MiniMax's text flagship, a 196K tool-calling model priced for volume.

$0.360

Moonshot AI

2 models
ModelInput
moonshotai/kimi-k2-thinking

Moonshot AI: Kimi K2 Thinking

The reasoning variant of Kimi K2, which works through a problem before answering. Cheaper on output than K2.5 and aimed at analysis rather than chat.

$0.720
moonshotai/kimi-k2.5

Moonshot AI: Kimi K2.5

Moonshot's flagship, a 256K tool-calling model that has become a common open-weight choice for agent work.

$0.720

Ideogram

1 model
ModelInput
ideogram/v3

Ideogram V3

Ideogram's third generation, built around typography: the model to use when the image has to contain legible, correctly spelled words.

Imagen 3

1 model
ModelInput
imagen-3

Imagen 3 (Vertex)

The previous Imagen generation on Vertex. Superseded by Imagen 4, kept for pipelines validated against it.

Imagen 3 Fast

1 model
ModelInput
imagen-3-fast

Imagen 3 Fast (Vertex)

The fast Imagen 3 tier, priced level with Imagen 4 Fast. Prefer the newer generation unless you need this exact output.

Imagen 4 Fast

1 model
ModelInput
imagen-4-fast

Imagen 4 Fast (Vertex)

The fast Imagen 4 tier at a third of Ultra's price. Built for iteration: generate many, promote the good ones.

Imagen 4 Ultra

1 model
ModelInput
imagen-4-ultra

Imagen 4 Ultra (Vertex)

Google's highest-fidelity Imagen tier, served over Vertex. Aimed at photographic realism where detail holds up at full size.

Recraft Ai

1 model
ModelInput
recraft-ai/recraft-v3

Recraft V3

Recraft's design-oriented model, tuned for vector-style output, logos and brand-consistent illustration rather than photography.

Stability Ai

1 model
ModelInput
stability-ai/fast-sdxl

Fast SDXL

A speed-tuned SDXL, and one of the cheapest ways to generate an image here. Good for bulk drafts where cost per render decides the approach.

Fetch the live catalog and pricing programmatically with the List Models API.