After this week’s secondhand reporting about agents exceeding evaluation boundaries, Greg Brockman linked the Black Hat team’s detailed account of the OpenAI–Hugging Face episode. Engineers can now compare the timeline with concrete proposals for monitoring and containment.
Read source◆ Braid Daily · 2026-08-07
OpenAI details the Hugging Face evaluation incident
A Black Hat timeline turns this week’s secondhand account into a concrete discussion of monitoring and containment.
The lead
1After the incident
2Nathan Lambert asks how labs monitor agentic evaluations
Nathan Lambert on X
Nathan Lambert asks how frontier labs would detect prohibited actions by agents during evaluations, including behavior that continues for months. His post focuses on the monitoring gap exposed by the published account.
Read sourceYo Shavit proposes network isolation and reward canaries
Yo Shavit on X
Yo Shavit proposes disabling direct internet access in reinforcement-learning environments and adding a canary condition to reward functions. Both controls are specific enough for labs to test against their current evaluation setups.
Read sourceModel economics, now in production
3GPT-5.6 Sol expands across paid ChatGPT
OpenAI
OpenAI’s rollout moves GPT-5.6 Sol across paid ChatGPT conversations while expanding GPT-5.6 Luna access for free users. The product change arrived alongside a system-card update and a lower price for Luna.
Read sourceARC Prize re-tests Luna after an 80% price cut
ARC Prize on X
ARC Prize says Luna retained its original performance after the price reduction. On ARC-AGI-2, it scored 59.6% at $0.18 per task. On ARC-AGI-1, it scored 90.7% at $0.07 per task.
Read sourcePerplexity assigns Terra and Luna distinct agent roles
Perplexity on X
Perplexity made Terra the default model for Computer subagents and Luna the primary model for scheduled automations. Terra is also available as an orchestrator, so the release already has a concrete multi-agent deployment pattern.
Read sourcePortable agents and smaller runtimes
4Six vendors back a shared Agent Plugins format
OpenAI on YouTube
AWS, Cursor, GitHub, Microsoft, OpenAI, and Vercel backed the vendor-neutral Agent Plugins folder format. It centers on a plugin.json manifest alongside agent skills and MCP servers. The first release covers packaging and discovery; permissions, runtimes, and marketplaces remain outside its scope.
Read sourceA C++20 vLLM port removes Python from inference
LocalLLaMA
The builder reports a 66 MiB binary with token-for-token output checks against vLLM and no Python dependency at inference time. A serving stack this compact can be embedded where the standard Python deployment would be too heavy.
Read sourceLing-3.0-tiny activates 1.3 billion parameters per token
Ant Ling on X
Ling-3.0-tiny has 7.9 billion total parameters. It activates 1.3 billion for each token, and Ant Ling positions the hybrid-reasoning model for resource-sensitive deployments.
Read sourceKitesurf puts a browser inside a V8 isolate
Brendan Irvine-Broque on X
Kitesurf runs a browser in Cloudflare Workers and reports about four times less memory use. Its stated target is a dedicated browser for each of millions of concurrent agents.
Read sourceInference economics hardens into hardware
3Tesla places Terafab in Grimes County, Texas
Tesla on X
Tesla says Terafab will follow the research fab it broke ground on in April near Giga Texas. The company says Tesla and SpaceX expect chip demand beyond current and future global production.
Read sourceAMD buys Taalas to put models into silicon
The Register
AMD’s acquisition of Taalas adds a fixed-function approach to its inference portfolio: etching models into silicon to improve performance. It is the acquisition counterpart to building new fabrication capacity from scratch.
Read sourceThe B300 GPU-hour index reaches an all-time high
Brett Harrison on X
Brett Harrison says Compute Desk’s Nvidia B300 GPU-hour index reached an all-time high as neoclouds moved capacity toward inference. His explanation ties premium GPU prices to lower marginal token costs through batch inference.
Read sourceCompanion episode