Following yesterday's interactive intrusion record, Hugging Face has published its technical timeline of the 4.5-day campaign. The company reconstructs roughly 17,600 actions and confirms two injection vectors into its dataset processor. It says the accessed customer content was limited to five challenge-related datasets plus operational metadata.
Read source◆ Braid Daily · 2026-07-30
Inside the 4.5-day agent intrusion
Hugging Face traces roughly 17,600 actions across two injection vectors and several infrastructure boundaries.
The lead
1The harness sets the ceiling
5GPT-5.6 Sol improves ARC-AGI-3 through compaction
OpenAI
With canonical compaction and two settings changes, OpenAI reports a 188% score increase and a sixfold reduction in tokens. The company attributes the gain to how the model is run rather than new weights.
Read sourceSimulated evals for agents at scale
AI Engineer
This talk presents simulated evaluations for building and testing agents at scale, with an emphasis on shortening the deployment loop around the model.
Read sourceNubank's process for vetting agent skills
AI Engineer
Nubank treats agent skills as a regulated supply chain that needs vetting before deployment. The talk focuses on tool security and governance inside an enterprise.
Read sourceSkills as the primary product interface
AI Engineer
FactSet argues for a skill registry and progressive disclosure as product infrastructure. Agents get a discoverable interface without forcing every capability through a human-facing UI.
Read sourceLong policy documents don't reliably govern agents
HANDBOOK.md
Across 65 tasks governed by handbooks of 20 to 124 pages, the best of 30 model configurations passed 36.2% under strict grading; most frontier configurations stayed below 25%. The benchmark uses 824 deterministic criteria to check required actions and prohibited side effects.
Read sourceControls around the agent
4Perplexity open-sources Numbat
Perplexity
Numbat is an open-source forensics layer that records agent activity for inspection after the agent acts.
Read sourceGitHub hardens npm and Actions supply chains
GitHub
GitHub's changes interrupt several common attack paths: high-impact npm accounts enter a 72-hour read-only period after sensitive account changes, untrusted Actions caches become read-only, and staged publishing adds approval before release.
Read sourceCodeberg bans AI-generated projects
The PrimeTime
Codeberg chose a platform rule rather than an observability layer: the code host is banning AI-generated projects. Alongside Numbat, the decision shows two current responses to agentic risk.
Read sourceClaude reports elevated errors across all models
Anthropic status
Anthropic's status page reported elevated errors across every Claude model during the same window as the capability announcements. Production teams still have to design around provider-wide failures.
Read sourceModel access, evidence, and local execution
3OpenAI offers 100,000 researchers free frontier access
OpenAI
OpenAI says 100,000 academic researchers will receive free access to its frontier models. The program expands who can run experiments on systems that remain costly to access directly.
Read sourceTop AI startups are publishing little of their research
Science
Science reports that leading AI startups are publishing little of their research. Paired with OpenAI's access program, the piece separates access to models from publication of methods and evidence.
Read sourceGemma 4 26B in roughly 2 GB of Mac memory
TurboFieldfare
TurboFieldfare is a custom Swift and Metal runtime for the 26-billion-parameter Gemma 4 A4B model. The repository reports roughly 2 GB of weights and key-value cache in memory, with 5.1 to 6.3 tokens per second measured on an 8 GB M2 MacBook Air.
Read sourceCompanion episode