Skip to content
Airating AIrating
Leaderboard
</> Best for code SWE-Bench, HumanEval, LiveCodeBench ◎ Logic & reasoning GPQA, ARC, complex tasks $ Cheapest models Lowest price per token ⚡ Fastest models Maximum throughput ⬚ Large context 128K to 10M tokens 🏆 All collections 10 categories ☰ All models Full catalog — search and filter the entire database
Quiz Compare Timeline Providers Wiki

Compare Models

Pick 2 models of the same type to compare

Same type only — showing

Model A
No models match your filters.
VS
Model B
No models match your filters.
Compare now Browse rankings

Popular comparisons

G Gemini 2.5 Pro 53.4 VS G Gemini 3.1 Pro Preview 86.7 N Nemotron 3 Ultra 550B A55B 81.1 VS O Openrouter z Ai GLM-4.7 74.8 C Claude Mythos 5 98.7 VS G GPT-5.5 Pro 100.0 H Hy3 77.9 VS Z Z.ai: GLM-5.2 91.5 M MiniMax M2.1 79.8 VS Z Z.ai: GLM-5.2 91.5 G Grok-4.1 92.8 VS G GPT-5.5 Pro 100.0 L Llama 3.2 90B Vision 67.0 VS Q Qwen-VL Max A Anthropic Claude Fable 5 92.7 VS C Claude Opus 4.7 Thinking M Mistral Small 3 VS Q Qwen 2.5 72B A Anthropic Claude Fable 5 92.7 VS C Claude Opus 4.6 Thinking 45.2 G GPT-5.3 Codex 92.9 VS G GPT-5.4 Pro 94.8 G GPT-4 Turbo 36.0 VS O o3 Mini 64.8 M MiniMax M2.7 52.8 VS G GPT-5.2 Codex 84.9 M MiniMax M2 74.1 VS Z Z.ai: GLM-5.2 91.5 C Claude Opus 4.6 Thinking 45.2 VS Z Z.ai: GLM-5.2 91.5

Data-driven AI model ratings. Updated regularly.

All models Rating Methodology Wiki Timeline About Airating Providers Privacy Policy Terms of Service Disclaimer