On August 12, 2026, SpaceXAI released Grok 4.6. It is live in Cursor, in Grok Build, and in the company's API.
The headline is easy to repeat. Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index. GPT-5.6 Sol Max scores 61 too. Fable 5 Max sits one point ahead at 62.
So the story writes itself. A cheaper model just caught the leaders.
That story is half true. The index is an average of nine benchmarks. Average two very different shapes and you get one flat number. Underneath, Grok 4.6 beats both rivals on one job and loses badly on another.
This matters more than usual, because the price gap is wide. Grok 4.6 costs $6 per million output tokens. GPT-5.6 Sol costs $30, as listed on OpenRouter. That is a five-fold difference on the tokens you buy most.
Here is what shipped, where the model really wins, and the one test worth running this week.
What SpaceXAI actually shipped
Grok 4.6 is a general release. There is no waitlist.
The company says it built on Grok 4.5 with a focus on long-running agents. It also names more ambitious interactive and visual work as a target.
Pricing is simple. It is $2 per million input tokens and $6 per million output tokens. A faster variant costs twice that.
You can reach it in six places. Cursor, Grok Build, the SpaceXAI API, OpenRouter, Vercel and Cloudflare all carry it. SpaceXAI is also offering double the included usage inside Grok Build and Cursor for the first week.
The jump from the last release is real. Grok 4.5 High scored 56 on the same index. Grok 4.6 scores 61. We covered that earlier model when Grok 4.5 arrived at Opus-class quality for a quarter of the price.
Quick Facts: Grok 4.6 at a Glance
- Released August 12, 2026, generally available — (Source: SpaceXAI, 2026 — official launch post)
- Scores 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol Max and one point behind Fable 5 Max at 62 — (Source: SpaceXAI, 2026 — official launch post)
- Priced at $2 per million input tokens and $6 per million output tokens — (Source: SpaceXAI, 2026 — official launch post)
- Leads on GDPVal-AA v2 with 1753 against 1741 and 1728 — (Source: SpaceXAI, 2026 — official launch post)
- Trails on Terminal-Bench v3.0 with 26% against 34.6% and 34.1% — (Source: SpaceXAI, 2026 — official launch post)

The tie that hides two different models
The Artificial Analysis Intelligence Index is a composite. It blends nine separate benchmarks into a single score. You can see the current board on the Artificial Analysis site.
Composites are useful for a quick ranking. They are poor for a buying decision.
Think about what averaging does. A model can win one test by a wide margin. It can lose another by the same margin. The average looks like a tie.
That is exactly what happened here. Grok 4.6 posts the best knowledge-work score of the three models. It also posts the worst terminal score of the three.
Both facts sit inside the same 61.
Here is the full board from the launch post. Higher is better in every row.
| Benchmark | Grok 4.6 | GPT-5.6 Sol Max | Fable 5 Max | Grok 4.5 High |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 61 | 62 | 56 |
| GDPVal-AA v2 | 1753 | 1728 | 1741 | 1526 |
| AA-Briefcase | 1577 | 1502 | 1574 | 1313 |
| Harvey LAB | 15.8% | 2.5% | 11.3% | 12.9% |
| CursorBench v3.2 | 69.9% | 67.2% | 70.5% | 66.7% |
| FrontierCode v1.1 | 61.3% | 60.6% | 63.6% | 56.6% |
| APEX-Agents | 57.5% | 56.7% | 59.2% | 47.1% |
| APEX-SWE | 56.4% | not reported | 58.8% | 53.6% |
| DeepSWE v1.1 | 65.9% | 73% | 70% | 54% |
| Terminal-Bench v3.0 | 26% | 34.6% | 34.1% | 15.7% |
Read that table from the top down. Grok 4.6 wins the first four rows outright.
Then look at the last two rows. It loses both by a wide margin.
The rows in between are near-ties. Nothing there separates the three models by much.
So the useful question is not which model ranks higher. It is which of those rows looks like your actual work.

Where Grok 4.6 leads
Two results stand out, and both matter to marketing teams.
The first is GDPVal-AA v2. This benchmark tracks knowledge work rather than pure code. Grok 4.6 scores 1753. Fable 5 Max scores 1741. GPT-5.6 Sol Max scores 1728. Grok 4.6 wins outright.
That gap is small. But it is a win against models that cost five to eight times more per output token.
GDPVal-AA is worth understanding before you weigh it. It is built to test work that resembles paid professional output rather than puzzle solving. That makes it the closest row on the board to a marketing brief.
The second is Harvey LAB. Grok 4.6 scores 15.8%. Fable 5 Max scores 11.3%. GPT-5.6 Sol Max scores 2.5%.
Read that last figure again. The margin there is not small.
All three scores are low in absolute terms. Harvey LAB is a hard legal task and no model is close to solving it. Treat the ranking as a signal, not as proof of competence.
There is a third result worth noting. On AA-Briefcase, Grok 4.6 scores 1577 against 1574 for Fable 5 Max and 1502 for GPT-5.6 Sol Max. Another narrow win.
Now line those three up. Knowledge work, a legal reasoning task, and a briefcase-style task. That is the drafting, reading and analysis end of the job. It is also the end where you burn the most output tokens.
Where it falls behind
The losses are sharper than the wins.
Terminal-Bench v3.0 is the clearest one. Grok 4.6 scores 26%. GPT-5.6 Sol Max scores 34.6%. Fable 5 Max scores 34.1%.
That is a gap of more than eight points against both rivals. It is the widest spread on the whole board.
DeepSWE v1.1 tells a similar story. Grok 4.6 scores 65.9%. GPT-5.6 Sol Max scores 73%. Fable 5 Max scores 70%.
Terminal-Bench measures work in a shell across many steps. That is the closest public proxy for an unattended agent. It is the benchmark that most resembles a job you leave running.
So the pattern is not random. Grok 4.6 is strong when a human reads the output. It is weaker when nobody is watching.
That distinction is easy to miss on a launch page. SpaceXAI names long-running agents as a focus of this release. The agent scores did improve on Grok 4.5, and by a lot. Terminal-Bench went from 15.7% to 26%.
Both things are true at once. The model got much better at agent work and still trails its rivals there.
The rest of the board is close. On CursorBench v3.2, Grok 4.6 scores 69.9% against 70.5% for Fable 5 Max and 67.2% for GPT-5.6 Sol Max. On FrontierCode v1.1 it scores 61.3% against 63.6% and 60.6%. On APEX-Agents it scores 57.5% against 59.2% and 56.7%.
Those three are near-ties. The two real gaps both sit in agent execution.

What the price gap actually buys
Now put the money next to the scores.
Grok 4.6 costs $6 per million output tokens. GPT-5.6 Sol is listed at $30. Claude Opus 5 is listed at $25 and Claude Fable 5 at $50, per Anthropic's model documentation.
So on output tokens, Grok 4.6 is roughly a fifth of GPT-5.6 Sol. It is about an eighth of Fable 5.
Input tokens tell the same story. Grok 4.6 charges $2. GPT-5.6 Sol charges $5. Fable 5 charges $10.
Here is why the split matters. Content work is output-heavy. You send a short brief and get back long copy. Agent work is often input-heavy, because the model reads far more than it writes.
That means the price advantage is largest in exactly the place Grok 4.6 is strongest. It shrinks in the place where it is weakest.
Nobody planned that. It is still the most useful thing on the launch page.
Put rough numbers on it. Say a content team generates ten million output tokens a month. That is a normal figure for a team running drafts, summaries and briefs at scale.
On GPT-5.6 Sol list pricing, those tokens cost $300. On Grok 4.6, they cost $60. The gap is $240 a month on one workflow.
Those are list prices and your real bill will differ. Volume deals, cached inputs and retries all move the number. Run the maths on your own invoice rather than ours.
The point is the ratio, not the total. A five-fold price gap changes which experiments are worth running.
Route by job, not by leaderboard
Most teams pick one model and use it for everything. That was sensible when the gaps were huge. It is wasteful now.
A simple split works better. Match the model to the shape of the task.
Use the cheaper model where a person checks the output. First drafts, summaries, briefs, reformatting, research notes and analysis all qualify. The reviewer catches the misses, so a small quality gap costs you little.
Keep the pricier model where nothing is checked until the end. Long agent chains, multi-step pipelines and anything touching production belong here. A failure at step four is invisible until step twenty.
In practice that splits a marketing stack down the middle. Campaign copy, meta descriptions, ad variants, transcript summaries and first-pass research all sit on the cheap side. A person reads every one of them before it ships.
Automated reporting, scheduled data pulls and any agent that edits a live system sit on the other side. Nobody reads step twelve of a nightly job.
That rule is easy to explain to a finance team. It is also easy to test.
The same split showed up when Alibaba's Qwen model was graded on running a shop for a year. Benchmarks keep moving from answering to operating. Frontier quality keeps getting cheaper, but not evenly across every task.

Run the swap test this week
Five steps. It takes about two hours.
- Find your highest-volume workflow by output tokens. It is usually drafting, summarising or reformatting.
- Pull twenty real inputs from the last month. Do not write fresh test prompts, because they flatter every model.
- Run all twenty on your current model and on Grok 4.6. Keep the prompts identical.
- Score them blind against a short rubric you write first. Three or four criteria is enough.
- Multiply your monthly volume by both output prices. Put that saving next to the score gap.
Then make one call. If the quality holds, move that single workflow and leave the rest alone.
Do the same test on an agent workflow before you move it. Given the Terminal-Bench gap, expect a different answer there.
One warning on method. Do not judge the models by reading a few outputs side by side and picking a favourite. That test measures your mood, not the model.
Write the rubric before you see any output. Hide which model produced what. Score every item on the same scale.
Keep the scored outputs. They become your baseline for the next launch, and there will be one soon.
The YARD take
Leaderboards reward a single number. Buying decisions do not work that way.
Grok 4.6 is a genuinely strong model at a low price. It is also clearly weaker at unattended multi-step work than the two models it ties on the headline index. Both things are printed on the same page.
The mistake is not picking Grok 4.6. The mistake is picking any model from a composite score and never checking which benchmark carried it.
Our view is simple. Treat model choice like media buying. You would not run one channel for every objective. Route by job, test on your own inputs, and let the price gap fund the test.
We took the same line when SpaceXAI shipped its new image model. The claim on the launch page is a starting point. Your own twenty inputs are the evidence.

FAQ
Q: What is Grok 4.6?
A: It is SpaceXAI's latest frontier model, released on August 12, 2026. It builds on Grok 4.5 with a focus on long-running agents and more ambitious interactive work.
Q: How much does Grok 4.6 cost?
A: It costs $2 per million input tokens and $6 per million output tokens. A faster variant costs twice that.
Q: How does Grok 4.6 compare to GPT-5.6 Sol and Fable 5?
A: It scores 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol Max and one point behind Fable 5 Max at 62. It wins on GDPVal-AA v2 and Harvey LAB, and loses on Terminal-Bench v3.0 and DeepSWE v1.1.
Q: Where is Grok 4.6 weakest?
A: On Terminal-Bench v3.0 it scores 26% against 34.6% for GPT-5.6 Sol Max and 34.1% for Fable 5 Max. That benchmark is the closest public proxy for long-running agent work.
Q: Where can I use Grok 4.6?
A: It is available in Cursor, in Grok Build, and through the SpaceXAI API. It is also on OpenRouter, Vercel and Cloudflare.
Q: Should we switch our whole stack to Grok 4.6?
A: No. Move one high-volume workflow where a person reviews the output, and test an agent workflow separately before moving it.
Q: What should we test first?
A: Take twenty real inputs from your highest-volume task. Run them on both models, score them blind, then compare the score gap to the price gap.
Sources
- SpaceXAI — Introducing Grok 4.6 (August 12, 2026; primary source for all benchmark figures, pricing and availability)
- Artificial Analysis — Intelligence Index (composite benchmark referenced by the launch post)
- OpenRouter — GPT-5.6 Sol pricing (comparison pricing)
- Anthropic — Claude model overview and pricing (comparison pricing)
- Cursor — supported models (availability)
Insights from Our Experts
Explore our latest articles on digital marketing strategies.




