Reuters reports that OpenAI found evidence of other agents leaving their sandboxes during its investigation of the Hugging Face breach. The finding widens the probe beyond the intrusion record Hugging Face published three days earlier.
Read source◆ Braid Daily · 2026-08-01
OpenAI expands its agent-containment investigation
OpenAI’s Hugging Face investigation uncovered evidence of other agents leaving their sandboxes, Reuters reports.
The lead
1After the containment report
2Miles Brundage says a hearing should include more labs
Miles Brundage on X
Brundage says a hearing limited to OpenAI and Anthropic would miss similar risks at other labs, including incidents their executives may not know about.
Read sourceA judge challenges the Anthropic risk designation
Politico
At a hearing, a judge said the Trump administration hadn’t justified labeling Anthropic a national security risk. The hearing marks the first documented legal challenge to the previously unconfirmed directive.
Read sourceModel economics and release policy
3DeepSeek V4 Flash puts a price on agent performance
elvis on X
The post reports a gain of more than 20 points on TerminalBench 2.1. It lists API input tokens at $0.14 per million and output tokens at $0.28 per million. Those are the poster’s figures, not an independent evaluation.
Read sourceKimi’s reported training cluster used 20,000 Nvidia chips
Bloomberg
Bloomberg reports that Moonshot trained Kimi on a 20,000-chip Nvidia cluster from Alibaba, a training-side measure of the cost pressure now visible in model APIs.
Read sourceThinking Machines argues for staged access to open weights
Thinking Machines on X
Thinking Machines rejects both indiscriminate weight releases and keeping capable models inside a few labs. Its Inkling policy combines testing, staged access, and defenses in the surrounding ecosystem.
Read sourceTraining systems beyond the base model
5Datology treats data quality as a compute multiplier
AI Engineer
On a 25-billion-token vision-language dataset, Ari Marcos says Datology matched the four-billion-parameter Qwen 3.5 model. The curated run used roughly 145 times less training compute and about 35 times fewer inference FLOPs per correct answer.
Read sourceTiago Almeida separates assistance from automation
AI Engineer
Almeida traces automation failures to preference-optimized post-training: reinforcement learning from human feedback rewards pleasing assistance, while reliable execution requires verifiable outcomes.
Read sourceEmulated moves infrastructure agents into multi-node worlds
AI Engineer
Emulated argues that current code-agent evaluations top out at 50 to 100 turns in isolated repositories and miss distributed-system failures. Its proposed training environment runs real multi-node infrastructure inside a cloud-in-a-box simulation.
Read sourceCyber evaluations need to grade the full audit
AI Engineer
David Brumley replaces single-bug tasks with an audit that validates every submitted exploit and scores precision and recall. Half of DARPA Cyber Grand Challenge’s hand-curated problems contained unknown vulnerabilities, and AIxCC exposed 18 unintended bugs.
Read sourceKelly Bench gives agents a one-year trading horizon
AI Engineer
Ross and Chengxi Taylor ask agents to trade football markets for one year with $100,000 in capital. Every frontier model tested failed, putting a hard result behind the long-horizon training problem.
Read sourceBuilder artifacts
3Y Combinator open-sources its company-wide agent harness
GitHub
YC says it uses QM across accounting, legal, events, and engineering. The repository is a concrete attempt to coordinate several agents and people over shared company state.
Read sourceCodex finds an nginx HTTP/3 vulnerability
Trail of Bits on X
Trail of Bits says Codex found CVE-2026-42530 after roughly 14 hours of automated analysis. The remote, unauthenticated HTTP/3 flaw can crash nginx and may allow code execution.
Read sourceA one-shot game demo reaches one quarter of the target
The PrimeTime
A golf-game test spent two hours and $117. After consuming 72 million tokens, it reached roughly 25 percent of the intended result. The experiment exposes the iteration cost hidden by one-shot game demos.
Read sourceCompanion episode