← All guidesAt a glance
Launched July 9, 2025 by xAI. Grok 4 was trained with heavy reinforcement learning on reasoning, and its Heavy configuration was the first model to break 50% on Humanity's Last Exam - a 2,500-question, PhD-level benchmark spanning maths, physics, chemistry, linguistics, and engineering.
On the All-inclusive plan a Grok 4 reply costs 3 credits.
Benchmarks
- Humanity's Last Exam (text subset): 50.7% - first past the 50% mark.
- ARC-AGI v2 (novel abstract reasoning): 15.9% - roughly double the next commercial model at release.
- USAMO 2025 maths-proof benchmark: 61.9% in the Heavy configuration.
Best at on Magnus
Hard reasoning with attitude: research questions, devil's-advocate takes, and debate rounds against a Claude agent. Its direct, slightly irreverent style makes it a favourite for a challenger persona on the agent bar.
- Grok covers chat and documents; integration tools (Gmail, Calendar, Slack actions) stay with Claude.