News Radar RSS

Grok 4.6 matched Fable 5 Max at an 85% discount. Downloadable models set that price.

The New Stack Cloud & Infrastructure Score 7/10

Summary

I’m Matt Burns, Chief Content Officer at Insight Media Group. Each week, I round up the most important AI developments, The post Grok 4.6 matched Fable 5 Max at an 85% discount. Downloadable models set that price. appeared first on The New Stack .

Original Text

I’m Matt Burns, Chief Content Officer at Insight Media Group. Each week, I round up the most important AI developments, explaining what they mean for people and organizations putting this technology to work. The thesis is simple: workers who learn to use AI will define the next era of their industries, and this newsletter is here to help you be one of them.

Grok 4.6 launched on Wednesday. Qwen 3.8-Max hit a few hours later. Then, on Thursday, DeepSeek V4-Pro followed. Three frontier models in about 24 hours, and all three pitched on cost because the capabilities are assumed.

It’s hard to argue against Elon Musk’s post on X: “Grok 4.6 is objectively #1 when considering intelligence, speed & cost”. Two years ago, frontier launches were about capability. Now they’re about price. Every benchmark existed to justify a cheaper bill.

Two of the three frontier launches this week came with downloadable weights, which is why the ceiling on what a closed lab can charge is increasingly set by companies giving their models away.

And that explains why Elon’s Cursor deal makes considerably more sense today than it did in June. When models converge on price, the money moves to whomever decides which model gets the job. Musk bought that.

Same model. Better training. Same bill.

SpaceXAI announced Grok 4.6 in two sentences. Frontier intelligence, the company said, and a “significant improvement over Grok 4.5 at the same price.” Artificial Analysis scored it 61 on its Intelligence Index, five points above 4.5 and even with GPT-5.6 Sol. On the GDPVal-AA leaderboard it went to the top at 1,753 Elo, just past Fable 5 Max at 1,741.

So where did the five points come from? Not a bigger model. As product analyst Aakash Gupta said, Grok 4.6 runs on the same 1.5 trillion parameters as 4.5 and every gain came out of post-training. Terminal-Bench jumped 66% in one release. The API price never moved, $2 in and $6 out per million tokens, because the same weights cost the same to serve. The improvement is, essentially, free.

The rest of the week explains that price.

Three frontier models in 24 hours, Aug. 11-13, 2026 Model

Weights you can download

Price (in / out)

Grok 4.6SpaceXAI

No

$2 / $6

DeepSeek-V4-Pro-0813DeepSeek

Yes

$0.66 / $1.98off-peak

Qwen3.8-MaxAlibaba

Yes2.4T params, 95B active

Self-host

Nemotron 3.5 LightningNvidia

Yes30B params, 3B active

Self-host

Claude Fable 5Anthropic, for reference

No

$10 / $50

Published API pricing per 1M tokens (input / output). DeepSeek rates are off-peak; peak runs double, and the new pricing takes effect Aug. 16, 2026. Fable 5 is shown for comparison and did not ship this week.

Our own Amanda Caswell went through the benchmarks SpaceXAI left out of those two sentences: 69.9% on CursorBench v3.2, a hair under Fable 5 Max at 70.5%, and behind the field on DeepSWE.

When Alibaba announced Qwen3.8-Max, Adrian Bridgwater caught the skepticism from consultant Jeff Brokaw, who called an open-weights promise with no release “the API business model wearing an open source jacket for the launch photo.” It was a fair take at the time. The weights are out now. DeepSeek followed with V4-Pro a day later, adding native support for OpenAI’s Response API and pricing table with peak and off-peak rates, off-peak running half of peak. Frontier intelligence now has nights and weekends minutes.

Against that, investor Gavin Baker did the math a buyer actually does. Grok 4.6 runs 80% cheaper on input tokens and 88% cheaper on output than Fable 5 Max. His verdict is two words: “Pareto dominant.”

If a developer’s fallback is a file on Hugging Face, a lab can’t simply charge more for a better model. It releases a better model at the old price and eats the difference.

When models become commodities, routing becomes the business

Box CEO Aaron Levie explained what happens next, and it’s worth your time to read it directly. But in short, cheaper models do not shrink AI budgets. They pay for work companies already wanted and could not afford, like agents scanning eerie codebase for security holes or reading every contract. He calls it Jevons paradox, and I think he’s right. The cheaper it gest, the more they buy. What follows from that, Levie wrote, is that more models tuned to different jobs and costs means “the more value there is in being at the layer that can route and optimize based on the task.”

Cheap models do not commoditize the stack evenly. The value slides toward whoever decides which model does what.

Nvidia shipped that exact product this week. TNS’ Frederic Lardinois covered Nemotron 3.5 Lightning, an open 30-billion-parameter model, alongside NeMo Switchyard, an open source router that sends each job to the model that suits it. Nvidia’s own numbers have a Switchyard setup pairing open models with Anthropic’s Opus 4.8 at about a third of the cost, with frontier accuracy intact.

Routers get more value the more interchangeable models become. A router turns model choice into a runtime decision instead of an architecture decision. Leaving a lab stops being a rewrite and starts being a config change, and a vendor you can swap easily has a hard time raising prices.

On Towards Data Science, Sara Nóbrega published the practical version of this, including the split she reports companies actually settling into: roughly 70% local small models, 20% mid-tier API, 10% frontier. She cites Nvidia research estimating that 40% to 70% of enterprise AI tasks already run on models under 10 billion parameters.

The obvious objection is that cheaper per token is not cheaper per job. It’s fair. Jessica Wachtel ran DeepSeek’s V4-Flash against V4-Pro for us and found the budget model reasoned better on the hardest task, then burned three times the tokens getting there. Final bill: Nine cents against ten. Sticker price tells you very little until you have run your own work through it.

Which brings me back to Cursor. SpaceX announced in June that it would buy the company in an all-stock deal valued at $60 billion. It’s worth noting this transaction has not yet closed. At the time it looked like a very expensive way to hire a coding-tools team. This week you could see what he bought. Michael Truell, the 25-year old co-founder and CEO of Cursor, helped launch Grok 4.6 and vouched for it in public writing that 4.6 “combines Opus-class intelligence and polish with very low cost and high speed.” The CEO behind Cursor is not publicly introducing Musk’s latest model.

Grok Bot launched the day before, pitched as AI teammates that “sign in to your tools, use them just like you do, and come back with finished work.” Engineer Kun Chen spent a day taking it apart and found Cursor underneath. Every user gets a cloud VM, the Mac and iOS apps are thin clients, and any real coding work gets handed off to Cursor’s cloud agents. Grok Bot, he wrote, “was clearly built and rebranded to Grok.”

So SpaceXAI now has the compute, the model, the agent harness and a few hundred million people to hand it to.

The compute is the part nobody else can go build. Musk put it inside SpaceX that compute and power would be the constraint, and Bloomberg reported that Anthropic alone now pays SpaceXAI $1.25 billion a month for it. This spring, SpaceXAI’s own models were using 11% of the computing power available to them, while Anthropic worked through a compute crunch.

Elon does not need to win the model race. He just needs to keep up. Likewise, the open-weight labs don’t have to win the frontier. They only have to set its price.

Here’s some advice that’s easier said than done: Pick one open-weight model and run real work through it. Not because Codex or Claude are disappearing. Because every discount in that pricing table exists because downloadable models give buyers another option (so try the options).

The post Grok 4.6 matched Fable 5 Max at an 85% discount. Downloadable models set that price. appeared first on The New Stack.

CloudInfrastructure

News Radar provides aggregated summaries. Full content and copyright remain with the original publisher.