Where Grok 4.6 wraps up a complicated task in roughly 53 steps, Claude Opus 5 takes around 103. That difference reveals more about what SpaceXAI has actually built than any top-line benchmark figure does.
Even with OpenAI, a notch below Anthropic
Grok 4.6 lands at 61 points on the Artificial Analysis Intelligence Index, drawing level with OpenAI’s GPT-5.6 Sol. The only models ahead of it are Anthropic's Claude Opus 5 at 63 and Claude Fable 5 at 62.
Compared with Grok 4.5, that’s a five-point improvement. The two points separating it from the index leader are slim enough that the index alone shouldn’t decide which model you use.
Agentic work is where those steps count
The model shines brightest on agentic tasks — the kind where it has to carry a multi-step workflow through on its own, with no human stepping in to redirect it.
GDPval-AA v2, a benchmark designed to gauge genuine knowledge work performed on a computer, places it second with an Elo of 1,753, trailing only Claude Opus 5.
Still, the step count deserves the closest attention. Cutting the steps for a task roughly in half cuts the tokens consumed along the way by roughly the same amount, and anybody who has watched an agent loop pile up charges knows that isn’t a rounding error.

Cost is the strongest case
Prices are unchanged at $2/$6 per million tokens. Set that against Claude Opus 5 at $5/$25 and GPT-5.6 Sol at $5/$30, and Grok 4.6 comes in more than 60 percent below both.
The gap hurts most on output tokens — precisely the ones agentic workloads churn out in bulk. When a model sits a couple of points down the benchmark index but costs a quarter as much on output, most production teams won’t hesitate.
Where to get it right now
You can use Grok 4.6 today via the API, Cursor, Grok Build and partners such as OpenRouter, Vercel and Cloudflare.
x.ai is also doubling usage quotas in Grok Build and Cursor for the opening week. Anyone running agents billed by the step should treat that as the window to find out whether 53 steps holds true on their own workload instead of inside a benchmark harness.

















STAY ALWAYS UP TO DATE