Benchmarks

Evolvent Frontier Index

Compare frontier models on each Evolvent benchmark using the metric reported by its original evaluation.

Anthropic

Claude Sonnet 4.6

Anthropic · avg@3

75.8%

80
60
40
20
0

Published evaluations

Benchmark Library

🥇
AnthropicClaude Sonnet 4.675.8
🥈
AnthropicClaude Opus 4.674.6
🥉
OpenAIGPT-5.4 High72.0
🥇
GoogleGemini 3.1 Pro Preview75.4
🥈
OpenAIGPT-563.3
🥉
AnthropicClaude Opus 4.661.3
🥇
AnthropicClaude Opus 532.5
🥈
AnthropicClaude Sonnet 526.3
🥉
GoogleGemini 3.1 Pro26.2
🥇
Moonshot AIKimi K3 + Kimi Code27.317%
🥈
OpenAICodex gpt-5.6-sol27.277%
🥉
AnthropicClaude Code Sonnet-523.441%

Stay posted on new benchmarks

Follow new evaluations, datasets, and reproducible agent research from Evolvent.

Benchmarks - Evolvent AI