Following yesterday’s outside look at harness regressions and machine-checked release gates, Anthropic’s Lance Martin explains how the company builds harnesses for long-horizon tasks. The talk focuses on practical choices for builders.
Read source◆ Braid Daily · 2026-07-23
Anthropic’s long-horizon agent harness
Lance Martin explains how Anthropic builds agent harnesses for work that must hold together over long tasks.
The lead
1Agent systems beyond the model
4NexForge generates training data for agents
arXiv cs.AI
NexForge generates agent training data and reports results that surpass proprietary models. It puts data generation beside the harness as another system-level source of capability.
Read sourceA benchmark built around payment integration
arXiv cs.AI
Alipay-PIBench tests coding agents on complex software work such as payment integration. It targets tasks that broad coding scores can miss.
Read sourceVerbatim chunks beat extracted memory artifacts
arXiv cs.AI
In a controlled comparison, verbatim chunks beat extracted artifacts by 15.9 points on LoCoMo and 22 points on LongMemEval-S. The authors recommend keeping structured artifacts alongside source text instead of using them as a replacement.
Read sourceVocabulary dropout targets diversity collapse
arXiv cs.AI
This paper proposes vocabulary dropout to address diversity collapse in advanced model training and curriculum design. The mechanism targets builders working with self-play and agentic training systems.
Read sourceAgents at work
3How software engineers are adapting to AI
The Guardian
The Guardian reports on how software engineers are responding to AI-driven changes in their work. Its focus on collective action adds the workers’ perspective to recent layoff and restructuring coverage.
Read sourceRobot corrections that persist across sessions
arXiv cs.RO
PhysClaw-0 retains language corrections across sessions while collecting robot-training data. The paper reports gains in data efficiency and task success from that retained guidance.
Read sourceRouterVLA reuses smoke-test data to select policies
arXiv cs.RO
RouterVLA uses smoke-test data as supervision for choosing vision-language-action policies. The paper reports a 14.64 percentage-point improvement in policy routing across robot deployments.
Read sourceCompanion episode