Archive BRAID DAILY
Astra's ten-problem claim enters public review
Subscribe

Braid Daily · 2026-08-02

Astra's ten-problem claim enters public review

OpenAI has published manuscripts, reasoning walkthroughs, and Lean certificates; independent mathematical review comes next.

Luminous mathematical forms suspended in a dark research chamber.

The lead

1

OpenAI says an internal Astra model produced ten results across mathematics and theoretical computer science. The company has published manuscripts, reasoning walkthroughs, and Lean certificates; independent mathematical review is the next test.

Read source

Capability reports under review

2

Mario Zechner reports a six-year DRM break

Mario Zechner on X

Mario Zechner reports that Kimi K3 broke DRM that had resisted attempts for more than six years, while GPT-5.6 Sol exceeded that result in his security test. This is a developer's first-person report, not a benchmark.

Read source

Agents that act and persist

4

Agent evaluation has to score a trajectory

Cameron R. Wolfe on X

Cameron Wolfe contrasts a single model response with an agent loop that reasons, calls tools, observes results, and repeats. Evaluation therefore has to cover the action trajectory as well as the final output.

Read source

Policy and security

3

Model economics and local infrastructure

3

DeepSeek V4-Flash gets a workload-cost comparison

Chubby on X

Following yesterday's V4-Flash pricing item, Chubby attributes to Artificial Analysis a claim that the model completes the same benchmark tasks as Fable 5 at 105 times lower total cost. The comparison is about total task cost, not token price alone.

Read source

Wafer benchmarks Kimi K3 on AMD MI355X

Wafer

Wafer reports 952 tokens per second per MI355X node for Kimi K3 and claims better performance per dollar than B300 at its stated GPU-hour prices. The write-up also discloses a slower cold-prefill result on MI355X.

Read source

Companion episode

Ten Problems, No List

· 00:27:45