Following yesterday's Grok 4.6 piece, three separate harnesses now put it near the frontier at roughly a third of the cost. One roundup calculates about $0.84 per benchmark task. That is 32% below GPT-5.6/Sonnet and 73% below Fable 5, though xAI still hasn't published a full model card.
Read source◆ Braid Daily · 2026-08-14
Grok 4.6 puts price per task beside frontier performance
Three harnesses put Grok 4.6 near the frontier. One roundup prices it at about $0.84 per task, as OpenAI and Google push speed and price.
The lead
1
Price, speed, and benchmark results
4CursorBench 3.2 puts Grok 4.6 first
X Freeze on X
On CursorBench, Grok 4.6 takes the top position in an independent coding test. The result adds a developer-workflow benchmark to the broader price comparison.
Read sourceRareBench adds a second Grok 4.6 result
Daniel McKinnon on X
RareBench reports a top result for Grok 4.6 alongside a substantial cost advantage over Claude Opus. That combination helps teams compare successful task completion against spend.
Read sourceGPT-5.6 Sol Ultrafast makes latency the product
Cerebras
Cerebras details the system behind OpenAI's 14-times speed claim. In OpenAI's incident-response example, work that took one to two hours falls to ten to fifteen minutes, changing which investigations are practical during an outage.
Read sourceGemini 3.7 Flash cuts the cheap tier again
Google says Gemini 3.7 Flash costs half as much as 3.6 Flash and emphasizes coding and web-development work. The release extends the day's competition on model cost.
Read sourceOpen weights reopen the tuning path
3dots3-note pairs 280 billion parameters with Apache 2.0
dots studio on X
dots3-note uses a mixture-of-experts design with 280 billion parameters and multimodal input. It has a 512 thousand-token context window, and the Apache 2.0 license lets teams tune and deploy the weights themselves.
Read sourceGLM 5.3 focuses on coding and cyber capability
Z.ai
The GLM 5.3 release centers frontier coding and emergent cyber capabilities. It gives builders another substantial open model to evaluate against their own repositories and security tasks.
Read sourcePaul Graham sees tuning attention returning to open weights
Paul Graham on X
Paul Graham points to renewed interest in tuning open-weight models. His comment coincides with two large releases that give teams concrete weights and licenses to work with.
Read sourceAgent infrastructure gets explicit triggers and endpoints
3Scheduled agents turn recurring work into a deployment primitive
Harrison Chase on X
Harrison Chase announces scheduled agents for work that should start without a fresh human prompt. A scheduler makes recurring and unattended runs part of the product contract rather than an external workaround.
Read sourceVercel puts nine agent CLIs behind one gateway
Chris Tate on X
Vercel's AI Gateway connects nine coding-agent command lines to shared credentials and billing. Teams can change the agent client without rebuilding the account and payment layer each time.
Read sourcePerplexity exposes its browsing agent through one API
Perplexity Developers on X
Perplexity's Agent API packages web search and browsing behind one hosted endpoint. The release gives developers a managed route to current web data without assembling each browsing component themselves.
Read sourceEvals built around real work
2Rails tests eight models on 21 framework tasks
Ruby on Rails on X
Rails maintainers tested eight models on 21 atomic tasks covering bugs, security, and features. They ran each task three times, producing a practical template for maintainers who want an evaluation tied to their own framework.
Read sourceClaude improves a prime-distribution bound after 650 attempts
Two Minute Papers
After roughly 650 prompts, an unreleased Claude improved a related bound without proving the Riemann hypothesis. The run's first crucial result appeared after a 37-minute computation. Its formalized proof is machine-checkable, and the model treated the result as probably already known.
Read source“Too strong to be new,”
Companion episode
Same Score, Third of the Price
This week's launches are now being judged by the economics of using them: price per completed task, latency during production work, and the infrastructure for unattended agents. Tests run inside teams' own repositories and operational workflows will show whether today's results transfer.