GPT-5 has arrived — and it's pushing the limits in math, coding, reasoning, and multimodal understanding.
We've gathered its latest benchmark scores and lined them up against other top models like Claude 4 Sonnet, Claude 4.1 Opus, Gemini 2.5, and Grok-4.
Let's dive in.
Humanity's Last Exam
A benchmark designed to test complex, cross-domain reasoning at the highest difficulty levels.

GPT-5 sets a new standard here, outperforming most peers across all tested domains.
Math
AIME
Measures advanced competition-level mathematical reasoning.

GPT-5 pro scored a perfect 100%, while GPT-5 (no tools) achieved 94.6% — the highest score among models without tool assistance.
Closest competitors were Grok-4 at 91.7% and Gemini 2.5 at 83.0%, with Claude 4.1 Opus trailing at 78.0%.

FrontierMath
Evaluates expert-level mathematical reasoning across some of the most challenging, research-grade problems.

GPT-5 pro (python) reached 32.1%, a notable jump from the previous 27.4% by the ChatGPT agent with full tool access. Without tools, GPT-5 scored 13.5%, still ahead of earlier generations.








