Benchmarks
Evolvent Frontier Index
Compare frontier models on each Evolvent benchmark using the metric reported by its original evaluation.
Anthropic
Claude Sonnet 4.6
Anthropic · avg@3
75.8%
Published evaluations
Benchmark Library
🥇
AnthropicClaude Sonnet 4.675.8
🥈
AnthropicClaude Opus 4.674.6
🥉
OpenAIGPT-5.4 High72.0
🥇
GoogleGemini 3.1 Pro Preview75.4
🥈
OpenAIGPT-563.3
🥉
AnthropicClaude Opus 4.661.3
🥇
AnthropicClaude Opus 532.5
🥈
AnthropicClaude Sonnet 526.3
🥉
GoogleGemini 3.1 Pro26.2
🥇
Moonshot AIKimi K3 + Kimi Code27.317%
🥈
OpenAICodex gpt-5.6-sol27.277%
🥉
AnthropicClaude Code Sonnet-523.441%
Stay posted on new benchmarks
Follow new evaluations, datasets, and reproducible agent research from Evolvent.