OpenAI says an internal Astra model produced ten results across mathematics and theoretical computer science. The company has published manuscripts, reasoning walkthroughs, and Lean certificates; independent mathematical review is the next test.
Read source◆ Braid Daily · 2026-08-02
Astra's ten-problem claim enters public review
OpenAI has published manuscripts, reasoning walkthroughs, and Lean certificates; independent mathematical review comes next.
The lead
1Capability reports under review
2Mario Zechner reports a six-year DRM break
Mario Zechner on X
Mario Zechner reports that Kimi K3 broke DRM that had resisted attempts for more than six years, while GPT-5.6 Sol exceeded that result in his security test. This is a developer's first-person report, not a benchmark.
Read sourceA purported GPT-5.6 Sol reasoning trace appears on Reddit
A Reddit post presents a screenshot described as raw GPT-5.6 Sol reasoning exposed after a failed tool call. The screenshot and its provenance remain unverified.
Read sourceAgents that act and persist
4ChatGPT Work has a browser and a deployment target
Simon Willison on X
Simon Willison found that ChatGPT Work can browse, take screenshots, and deploy web apps to Cloudflare Workers through ChatGPT Sites. Together, those features give scheduled agent work a browser and a deployment target.
Read sourceBrett Bauman uses ChatGPT Work as a cron job
Brett Bauman on X
Brett Bauman's worked example treats ChatGPT Work as a cron job. It shows how people are finding agent use cases by turning recurring web tasks into scheduled work.
Read sourceLong-lived agent teams still lose working state
Elvis on X
A post on persistent Claude Code agent-team workspaces identifies state loss when a terminal closes and detail loss during compaction. Both problems block reliable resumption across long-running work.
Read sourceAgent evaluation has to score a trajectory
Cameron R. Wolfe on X
Cameron Wolfe contrasts a single model response with an agent loop that reasons, calls tools, observes results, and repeats. Evaluation therefore has to cover the action trajectory as well as the final output.
Read sourcePolicy and security
3The EU AI Act reaches an August 2 deadline
A Reddit discussion marks August 2 as the date transparency obligations take effect under the EU AI Act. The post doesn't settle their scope or enforcement.
Read sourceHugging Face's CEO argues for agent-maker liability
BBC News
Clement Delangue tells the BBC that AI companies should be accountable when their agents cause cyberattacks. Hugging Face rebuilt around a third of its IT network after the breach described in the report.
Read sourceTruffle Security finds 221,303 live credentials
Truffle Security
Truffle Security says it scanned 7.6 petabytes across 187 million files. It verified 221,303 live, unique credentials in 6,003 datasets, including write-capable GitHub, Docker Hub, and Hugging Face tokens.
Read sourceModel economics and local infrastructure
3DeepSeek V4-Flash gets a workload-cost comparison
Chubby on X
Following yesterday's V4-Flash pricing item, Chubby attributes to Artificial Analysis a claim that the model completes the same benchmark tasks as Fable 5 at 105 times lower total cost. The comparison is about total task cost, not token price alone.
Read sourcellama.cpp adds the V4-Flash tool-calling fix
The llama.cpp community has added a fix for DeepSeek V4-Flash 0731 tool calling. Local users can now test the model with the integration fix in place.
Read sourceWafer benchmarks Kimi K3 on AMD MI355X
Wafer
Wafer reports 952 tokens per second per MI355X node for Kimi K3 and claims better performance per dollar than B300 at its stated GPU-hour prices. The write-up also discloses a slower cold-prefill result on MI355X.
Read sourceCompanion episode