ApexApiApexApi
Sign In

AI Model Rankings

Measured success rate and latency for the models we serve, from real API calls made through ApexApi during health sweeps over the last 30 days. These are our own measured numbers, not vendor claims.

#ModelSuccessAvg latency
1Gemini 3.1 Flash Lite

google

100%1.8s
2Gemini 3.1 Flash Lite Preview

google

100%1.9s
3GPT-4.1 Mini

openai

100%2.4s
4Gemini 3.6 Flash

google

100%2.6s
5GPT-4.1

openai

100%2.7s
6Ministral 3 8B 2512

mistral

100%2.8s
7Ministral 3 3B 2512

mistral

100%2.8s
8Mistral Medium 3

mistral

100%2.8s
9Gemini 2.5 Flash Lite

google

100%2.9s
10GPT-OSS 20B

bedrock

100%2.9s
11Claude Haiku 4.5

anthropic

100%2.9s
12GPT-5.6 Luna

openai

100%3.0s
13Gemini 3.5 Flash Lite

google-vertex

100%3.0s
14GPT-5.5

openai

100%3.1s
15GPT-5.4 Mini

openai

100%3.2s
16GPT-4o

openai

100%3.2s
17GPT-5.1

openai

100%3.2s
18Mistral Small 4

mistral

100%3.3s
19Gemini 3 Flash Preview

google

100%3.3s
20Gemini 3.5 Flash

google

100%3.3s
21GLM 4.7 Flash

bedrock

100%3.3s
22o4 Mini

openai

100%3.3s
23GPT-4o

openai

100%3.3s
24GPT-4o

openai

100%3.3s
25o3

openai

100%3.4s
26Claude Haiku 4.5

anthropic-vertex

100%3.4s
27GPT-5.6 Terra

openai

100%3.4s
28GPT-4o-mini

openai

100%3.5s
29Claude Opus 4.5

anthropic-vertex

100%3.5s
30GPT-OSS 120B

bedrock

100%3.5s
31Claude Opus 4.8

anthropic-vertex

100%3.5s
32Kimi K2.5

bedrock

100%3.5s
33GPT-4o

openai

100%3.5s
34Nano Banana

google

100%3.5s
35GPT-5.4 Nano

openai

100%3.5s
36Ministral 3 8B 2512

bedrock

100%3.6s
37Mistral Medium 3.5

mistral

100%3.6s
38GPT-4o-mini

openai

100%3.6s
39Devstral 2 2512

bedrock

100%3.6s
40GPT-5.4

openai

100%3.6s
41GPT-5

openai

100%3.6s
42Gemma 4 26B A4B

google

100%3.6s
43Codestral 2508

mistral

100%3.7s
44Claude Opus 4.8

anthropic

100%3.7s
45GPT-5.2

openai

100%3.7s
46GPT-4.1 Nano

openai

100%3.7s
47Claude Sonnet 4.6

anthropic

100%3.7s
48Ministral 3 14B 2512

mistral

100%3.8s
49GPT-3.5 Turbo

openai

100%3.8s
50o3 Mini

openai

100%3.8s
51Ministral 3 3B 2512

bedrock

100%3.8s
52GLM 4.7

bedrock

100%3.9s
53Qwen3 VL Flash

alibaba

100%4.0s
54Ministral 3 14B 2512

bedrock

100%4.0s
55GPT-4

openai

100%4.1s
56Qwen3 VL Plus

alibaba

100%4.1s
57Qwen3.6 Flash

alibaba

100%4.2s
58o1

openai

100%4.2s
59Claude Opus 4.7

anthropic

100%4.2s
60Claude Opus 4.7

anthropic-vertex

100%4.2s
61Claude Opus 5

anthropic-vertex

100%4.4s
62Gemma 4 31B

google

100%4.5s
63Claude Opus 4.6

anthropic

100%4.5s
64DeepSeek V4 Flash

deepseek

100%4.5s
65GPT-5 Mini

openai

100%4.6s
66Claude Opus 4.5

anthropic

100%4.6s
67Gemini 3.1 Pro Preview Custom Tools

google

100%4.7s
68GPT-5.6 Sol

openai

100%4.7s
69Qwen3 Coder Plus

alibaba

100%4.7s
70GPT-4 Turbo

openai

100%4.9s
71DeepSeek V4 Pro

deepseek

100%4.9s
72Gemini 3.1 Pro Preview

google

100%4.9s
73Claude Sonnet 4.6

anthropic-vertex

100%5.0s
74GPT-5 Nano

openai

100%5.1s
75Qwen3.5 Flash

alibaba

100%5.3s
76Qwen3 Max

alibaba

100%5.5s
77GPT-3.5 Turbo 16k

openai

100%5.5s
78Claude Opus 4.6

anthropic-vertex

100%5.6s
79Gemini 2.5 Flash

google

100%5.8s
80Qwen3.7 Max

alibaba

100%6.1s
81GLM 5

bedrock

100%6.8s
82Qwen3.7 Plus

alibaba

100%8.0s
83Gemini 2.5 Pro

google

100%9.4s
84Kimi K2 Thinking

bedrock

100%9.4s
85MiniMax M2.5

bedrock

100%12.6s
86Devstral 2 2512

mistral

100%15.2s
87Qwen3.5 Plus

alibaba

100%87.0s
88GPT-4o Search Preview

openai

0%
89Mistral Large 3 2512

bedrock

0%
90GPT-4o-mini Search Preview

openai

0%
91Mistral Large 3 2512

mistral

0%

Latency is the average of successful calls and depends on prompt size, region, and upstream load. Success rate reflects our smoke tests, not a guarantee. Call any of these through one API and key at the model catalog.