Archive BRAID DAILY
Grok 4.6 puts price per task beside frontier performance
Subscribe

Braid Daily · 2026-08-14

Grok 4.6 puts price per task beside frontier performance

Three harnesses put Grok 4.6 near the frontier. One roundup prices it at about $0.84 per task, as OpenAI and Google push speed and price.

Three luminous compute streams cross a dark editorial surface, led by a fast signal-yellow path.

The lead

1

Following yesterday's Grok 4.6 piece, three separate harnesses now put it near the frontier at roughly a third of the cost. One roundup calculates about $0.84 per benchmark task. That is 32% below GPT-5.6/Sonnet and 73% below Fable 5, though xAI still hasn't published a full model card.

Read source
Grok 4.6 benchmark and cost signals across CursorBench, RareBench, Artificial Analysis, and per-task comparisons.
The day's Grok 4.6 reports separate capability signals from the price-per-task argument.

Price, speed, and benchmark results

4

RareBench adds a second Grok 4.6 result

Daniel McKinnon on X

RareBench reports a top result for Grok 4.6 alongside a substantial cost advantage over Claude Opus. That combination helps teams compare successful task completion against spend.

Read source

GPT-5.6 Sol Ultrafast makes latency the product

Cerebras

Cerebras details the system behind OpenAI's 14-times speed claim. In OpenAI's incident-response example, work that took one to two hours falls to ten to fifteen minutes, changing which investigations are practical during an outage.

Read source

Open weights reopen the tuning path

3

Agent infrastructure gets explicit triggers and endpoints

3

Evals built around real work

2

Rails tests eight models on 21 framework tasks

Ruby on Rails on X

Rails maintainers tested eight models on 21 atomic tasks covering bugs, security, and features. They ran each task three times, producing a practical template for maintainers who want an evaluation tied to their own framework.

Read source

Claude improves a prime-distribution bound after 650 attempts

Two Minute Papers

After roughly 650 prompts, an unreleased Claude improved a related bound without proving the Riemann hypothesis. The run's first crucial result appeared after a 37-minute computation. Its formalized proof is machine-checkable, and the model treated the result as probably already known.

“Too strong to be new,”

Read source

Companion episode

Same Score, Third of the Price

· 00:25:21

This week's launches are now being judged by the economics of using them: price per completed task, latency during production work, and the infrastructure for unattended agents. Tests run inside teams' own repositories and operational workflows will show whether today's results transfer.