◆ Dispatch 078 · 2026-07-06 GSV The Lease Asked for a Substation
The Lease Needed Power
“The AI buildout is being measured in leases, rack dates, grid queues, and rules about what synthetic characters are allowed to do.”
— Lenar Kess, today's narration
Today’s episode follows the AI buildout as it leaves the slide deck and runs into financing, rack design, grid queues, and product rules. The largest story is a set of infrastructure receipts: a Treasury warning, a reported Nvidia rack delay, Anthropic’s Kentucky lease, and a Scottish data-center project whose renewable-power promise does not appear to survive contact with local evidence.
- Techmeme’s NOTUS item on a draft Treasury report puts market-risk language next to the AI capex boom, with dotcom-era comparisons entering the official vocabulary.
- Techmeme’s CNBC item on Nvidia Kyber NVL144 tracks a reported 12-month-plus delay and cancellation of NVL72x2, which makes rack-scale interconnects part of the AI timeline.
- Techmeme’s CNBC item on Anthropic and TeraWulf gives the other side of the same story: a 20-year, $19 billion Kentucky data-center lease with about 400 megawatts of planned capacity.
- The Guardian’s Lanarkshire investigation shows how quickly a national AI growth-zone promise turns into land, planning, grid access, and renewable-power math.
- Techmeme’s China agent-rule item and Rest of World’s web-novel reporting show Chinese platforms narrowing what users can create with AI before regulators and readers force the issue harder.
- Armin Ronacher’s Claude tool-call critique, Fly.io’s agent sandbox post, and Harrison Chase’s harness note turn the builder segment toward model-harness contracts rather than demo polish.
- Forbes on Sam Altman’s global-referee proposal, Forbes on FTC bias disclosure, and The Guardian on the FCA Mills review give three different versions of AI governance: standards, disclosure, and financial-services supervision.
Chapters
- 00:00:04 Transcript
Sources
21 cited-
1
@hwchase17 (Harrison Chase)
X
This identifies a major shift in tooling/frameworks for AI agents ('agent harnesses' vs 'agent frameworks'), which is a primary builder artifact change and directly impacts development workflows.
x.com/hwchase17/status/2073793142739132625 →Details
- Context
- This identifies a major shift in tooling/frameworks for AI agents ('agent harnesses' vs 'agent frameworks'), which is a primary builder artifact change and directly impacts development workflows.
- Key points
- This identifies a major shift in tooling/frameworks for AI agents ('agent harnesses' vs 'agent frameworks'), which is a primary builder artifact change and directly impacts development workflows.
- Provenance
- Tweet · Primary source
-
2
Techmeme - Industry Adjacent (US)
Article
Directly addresses regulatory intervention (China's rules) impacting major Chinese AI players (ByteDance/Alibaba), specifically targeting agentic capabilities and human-AI interaction.
www.techmeme.com/260705/p9 →Details
- Context
- Directly addresses regulatory intervention (China's rules) impacting major Chinese AI players (ByteDance/Alibaba), specifically targeting agentic capabilities and human-AI interaction.
- Key points
- Directly addresses regulatory intervention (China's rules) impacting major Chinese AI players (ByteDance/Alibaba), specifically targeting agentic capabilities and human-AI interaction.
- Provenance
- Article · Supporting source
-
3
Building Agents That Don't Break Themselves — 4 pts · 0 comments
Article
The story addresses agentic tools and reliability in AI agents, a core focus area for senior builders interested in practical development workflows.
fly.io/blog/building-agents-that-dont-break… →Details
- Context
- The story addresses agentic tools and reliability in AI agents, a core focus area for senior builders interested in practical development workflows.
- Key points
- The story addresses agentic tools and reliability in AI agents, a core focus area for senior builders interested in practical development workflows.
- Provenance
- Article · Supporting source
-
4
r/LocalLLaMA: Llama-Server is Throwing Away Your Perfectly Good KV Caches, and How to Fix It - 0 pts · 0 comments
Article
Identifies a critical functional bug in a major local LLM serving tool (llama-server), directly impacting long-context efficiency and developer workflows.
www.reddit.com/r/LocalLLaMA/comments/1uohso… →Details
- Context
- Identifies a critical functional bug in a major local LLM serving tool (llama-server), directly impacting long-context efficiency and developer workflows.
- Key points
- Identifies a critical functional bug in a major local LLM serving tool (llama-server), directly impacting long-context efficiency and developer workflows.
- Provenance
- Article · Supporting source
-
5
The Guardian AI - Industry Adjacent (UK)
Article
Focuses on a major technical bottleneck (dextrous hands) for embodied AI/humanoid robots, linking it to geopolitical competition in China.
www.theguardian.com/technology/ng-interacti… →Details
- Context
- Focuses on a major technical bottleneck (dextrous hands) for embodied AI/humanoid robots, linking it to geopolitical competition in China.
- Key points
- Focuses on a major technical bottleneck (dextrous hands) for embodied AI/humanoid robots, linking it to geopolitical competition in China.
- Provenance
- Article · Supporting source
-
6
Forbes Innovation - Industry Adjacent (US)
Article
Discusses 'Agent Gateways' as a control plane for enterprise AI, signaling major architectural shifts in how companies deploy and govern AI.
www.forbes.com/sites/janakirammsv/2026/07/0… →Details
- Context
- Discusses 'Agent Gateways' as a control plane for enterprise AI, signaling major architectural shifts in how companies deploy and govern AI.
- Key points
- Discusses 'Agent Gateways' as a control plane for enterprise AI, signaling major architectural shifts in how companies deploy and govern AI.
- Provenance
- Article · Supporting source
-
7
Techmeme - Industry Adjacent (US)
Article
Reports on 'agentic ransomware,' linking AI capabilities (adaptivity/retries) to a major threat vector. This is a breaking security story with clear downstream consequence for enterprise infrastructure and risk manageme…
www.techmeme.com/260706/p1 →Details
- Context
- Reports on 'agentic ransomware,' linking AI capabilities (adaptivity/retries) to a major threat vector. This is a breaking security story with clear downstream consequence for enterprise infrastructure and risk management.
- Key points
- Reports on 'agentic ransomware,' linking AI capabilities (adaptivity/retries) to a major threat vector. This is a breaking security story with clear downstream consequence for enterprise infrastructure and risk management.
- Provenance
- Article · Supporting source
-
8
Techmeme - Industry Adjacent (US)
Article
Major product delay/cancellation (NVL144) for a key AI infrastructure component (Nvidia rack system). Directly impacts compute availability and industry timelines.
www.techmeme.com/260706/p3 →Details
- Context
- Major product delay/cancellation (NVL144) for a key AI infrastructure component (Nvidia rack system). Directly impacts compute availability and industry timelines.
- Key points
- Major product delay/cancellation (NVL144) for a key AI infrastructure component (Nvidia rack system). Directly impacts compute availability and industry timelines.
- Provenance
- Article · Supporting source
-
9
The Guardian Technology - Industry Adjacent (UK)
Article
Exposes a major infrastructure/power issue for a large AI datacenter project (£8.2bn). Directly relates to energy, power supply, and geopolitical feasibility of AI growth.
www.theguardian.com/technology/2026/jul/06/… →Details
- Context
- Exposes a major infrastructure/power issue for a large AI datacenter project (£8.2bn). Directly relates to energy, power supply, and geopolitical feasibility of AI growth.
- Key points
- Exposes a major infrastructure/power issue for a large AI datacenter project (£8.2bn). Directly relates to energy, power supply, and geopolitical feasibility of AI growth.
- Provenance
- Article · Supporting source
-
10
The Guardian AI - Industry Adjacent (UK)
Article
Examines government plans for massive AI datacentre zones (500MW+), questioning feasibility and reliability of infrastructure promises.
www.theguardian.com/technology/2026/jul/06/… →Details
- Context
- Examines government plans for massive AI datacentre zones (500MW+), questioning feasibility and reliability of infrastructure promises.
- Key points
- Examines government plans for massive AI datacentre zones (500MW+), questioning feasibility and reliability of infrastructure promises.
- Provenance
- Article · Supporting source
-
11
Forbes Innovation - Industry Adjacent (US)
Article
FTC policy action is a major regulatory intervention concerning LLM bias disclosure, directly impacting how AI models are built and deployed.
www.forbes.com/sites/lanceeliot/2026/07/06/… →Details
- Context
- FTC policy action is a major regulatory intervention concerning LLM bias disclosure, directly impacting how AI models are built and deployed.
- Key points
- FTC policy action is a major regulatory intervention concerning LLM bias disclosure, directly impacting how AI models are built and deployed.
- Provenance
- Article · Supporting source
-
12
Techmeme - Industry Adjacent (US)
Article
Directly addresses a core developer workflow (tool calling) and critiques major model releases (Claude Opus/Sonnet), providing actionable insight for builders.
www.techmeme.com/260706/p4 →Details
- Context
- Directly addresses a core developer workflow (tool calling) and critiques major model releases (Claude Opus/Sonnet), providing actionable insight for builders.
- Key points
- Directly addresses a core developer workflow (tool calling) and critiques major model releases (Claude Opus/Sonnet), providing actionable insight for builders.
- Provenance
- Article · Supporting source
-
13
Techmeme - Industry Adjacent (US)
Article
A draft US Treasury report warning about AI risks and comparing it to the dotcom crash is a major regulatory/financial signal that impacts market structure and capital allocation.
www.techmeme.com/260706/p5 →Details
- Context
- A draft US Treasury report warning about AI risks and comparing it to the dotcom crash is a major regulatory/financial signal that impacts market structure and capital allocation.
- Key points
- A draft US Treasury report warning about AI risks and comparing it to the dotcom crash is a major regulatory/financial signal that impacts market structure and capital allocation.
- Provenance
- Article · Supporting source
-
14
Rest of World Latest - Media Culture (GLOBAL)
Article
Major Chinese tech players (Tencent, ByteDance, Baidu) are implementing content controls on AI-generated content, signaling a regulatory/market shift in creative control and quality standards.
restofworld.org/2026/china-ai-web-novels →Details
- Context
- Major Chinese tech players (Tencent, ByteDance, Baidu) are implementing content controls on AI-generated content, signaling a regulatory/market shift in creative control and quality standards.
- Key points
- Major Chinese tech players (Tencent, ByteDance, Baidu) are implementing content controls on AI-generated content, signaling a regulatory/market shift in creative control and quality standards.
- Provenance
- Article · Supporting source
-
15
Forbes Innovation - Industry Adjacent (US)
Article
Altman calling for global standards and offering the US government equity is a major corporate/policy dynamic that directly impacts control and governance.
www.forbes.com/sites/anishasircar/2026/07/0… →Details
- Context
- Altman calling for global standards and offering the US government equity is a major corporate/policy dynamic that directly impacts control and governance.
- Key points
- Altman calling for global standards and offering the US government equity is a major corporate/policy dynamic that directly impacts control and governance.
- Provenance
- Article · Supporting source
-
16
Techmeme - Industry Adjacent (US)
Article
Covers corporate dynamics (SPAC IPO) and physical AI/robotics market structure, which is a key area of power struggle.
www.techmeme.com/260706/p11 →Details
- Context
- Covers corporate dynamics (SPAC IPO) and physical AI/robotics market structure, which is a key area of power struggle.
- Key points
- Covers corporate dynamics (SPAC IPO) and physical AI/robotics market structure, which is a key area of power struggle.
- Provenance
- Article · Supporting source
-
17
ChinaTalk - Policy Geopolitics (CN)
Article
Discusses physical-world AI (robotics) and potential geopolitical/labor shifts in China, directly impacting where intelligence is built.
www.chinatalk.media/p/the-robots-are-here →Details
- Context
- Discusses physical-world AI (robotics) and potential geopolitical/labor shifts in China, directly impacting where intelligence is built.
- Key points
- Discusses physical-world AI (robotics) and potential geopolitical/labor shifts in China, directly impacting where intelligence is built.
- Provenance
- Article · Supporting source
-
18
The Guardian Technology - Industry Adjacent (UK)
Article
A landmark FCA review detailing regulatory gaps in financial services (AI/cybercrime) is a major policy intervention and signals future industry constraints.
www.theguardian.com/business/2026/jul/06/bo… →Details
- Context
- A landmark FCA review detailing regulatory gaps in financial services (AI/cybercrime) is a major policy intervention and signals future industry constraints.
- Key points
- A landmark FCA review detailing regulatory gaps in financial services (AI/cybercrime) is a major policy intervention and signals future industry constraints.
- Provenance
- Article · Supporting source
-
19
The Guardian AI - Industry Adjacent (UK)
Article
Discusses AI-powered surveillance systems tracking public/private life, directly impacting policy, civil liberties, and government control.
www.theguardian.com/commentisfree/2026/jul/… →Details
- Context
- Discusses AI-powered surveillance systems tracking public/private life, directly impacting policy, civil liberties, and government control.
- Key points
- Discusses AI-powered surveillance systems tracking public/private life, directly impacting policy, civil liberties, and government control.
- Provenance
- Article · Supporting source
-
20
Techmeme - Industry Adjacent (US)
Article
Major infrastructure/capital allocation story (Anthropic, $19B lease, 400MW capacity). Directly relates to AI infrastructure and power struggles.
www.techmeme.com/260706/p16 →Details
- Context
- Major infrastructure/capital allocation story (Anthropic, $19B lease, 400MW capacity). Directly relates to AI infrastructure and power struggles.
- Key points
- Major infrastructure/capital allocation story (Anthropic, $19B lease, 400MW capacity). Directly relates to AI infrastructure and power struggles.
- Provenance
- Article · Supporting source
-
21
Better Models: Worse Tools
Article Armin Ronacher — Developer and creator of Flask, writing from direct debugging of Pi tool-call failures.
Turning on strict tool invocation eliminated it in my runs.
lucumr.pocoo.org/2026/7/4/better-models-wor… →Details
- Cited text
Turning on strict tool invocation eliminated it in my runs.
- Context
- It gives the agent segment a concrete model-harness failure rather than a generic reliability complaint.
- Key points
- Newer Claude models produced correct edit payloads but added extra keys inside Pi's nested edits array.
- The failure appeared after agentic history rather than in fresh single-turn prompts.
- Ronacher argues model post-training may be adapting strongly to Claude Code-like harness behavior.
- Provenance
- Article · Supporting source
Transcript
00:00:04 lenarTechmeme’s top AI infrastructure items this morning put four different facts next to each other: a draft Treasury report warning about AI-market risk, a reported Nvidia rack delay, Anthropic signing a huge Kentucky data-center lease, and a Guardian investigation into a Scottish AI project whose renewable-power promise doesn't appear to hold up. [pause] Those aren't four versions of one story. They are four places where the AI buildout meets a different test. A finance warning becomes credit discipline. A rack design meets the factory. A lease needs megawatts on a schedule. And a political promise about green data centers runs into land, grid access, planning approval, and cables in the ground.
00:00:50 damraThe order matters there. The fantasy version of the AI boom starts with models and then assumes everything around them catches up. Today’s source list starts with the physical world asking for receipts. Can the rack ship? Can the site draw power? Can the lease turn into usable capacity before the customer’s model roadmap needs it? Can the money wait that long without getting nervous?
00:01:14 lenarThe Treasury item is via Eric Katz at NOTUS, surfaced on Techmeme. The NOTUS report says a draft US Treasury Department document is set to warn about AI-market risks and compare parts of the current moment to the dotcom crash. That’s not a crash call. I wouldn’t treat it that way. It is a sign that the official financial vocabulary is catching up to the scale of the spend. Once the government starts writing about AI capex with dotcom-era comparisons, the conversation has moved from ‘can these companies raise enough money?’ to ‘what happens if a lot of these promises are capitalized before the returns arrive?’
00:01:51 damraThat creates a different kind of pressure than a product delay. A product delay hurts the roadmap. A financial-risk memo changes who has permission to be skeptical. You can have a CFO, a regulator, or a lender say, ‘show me the signed customers, show me the power contract, show me the delivery date,’ and that request no longer sounds like they don’t understand AI. It sounds like they read the same memo.
00:02:17 lenarThen you have the Nvidia item. CNBC, citing SemiAnalysis, reports that Nvidia’s Kyber NVL144 rack-scale system has slipped by more than 12 months to 2028 because of printed-circuit-board manufacturing issues, and that the NVL72x2 architecture has been canceled. The component that moved the date matters more than the date itself. This isn't a model waiting for another benchmark pass. It is the physical rack system around Rubin Ultra. Scale-up interconnects have to work, copper and co-packaged optics have to meet the design, and the midplane has to be manufacturable before that promised compute tier exists.
00:02:59 damraThe point that stuck with me is that a printed circuit board becomes a model-timeline event. If Kyber is delayed because the midplane is hard to manufacture, then the AI roadmap is no longer only about chip supply. It’s about whether the rack-scale architecture can move enough data across enough accelerators without becoming a science project. The model people can be completely right about demand. They can still end up waiting on electrical engineering paperwork until the schedule breaks.
00:03:27 lenarRight. And in the same morning window, Techmeme has CNBC reporting that Anthropic signed a 20-year, $19 billion lease with TeraWulf for a Kentucky data center expected to reach roughly 400 megawatts of capacity, with first power delivery in the second half of 2027. That isn't a vague cloud partnership. It is a long lease on a power-hungry physical site. The supplier is named, the state is named, the capacity figure is named, and the first-power window is part of the report.
00:03:59 damraIt also tells you why the Treasury-style skepticism and the Nvidia-style schedule risk belong in the same conversation, even if they aren't the same claim. Anthropic can have extraordinary demand. It can also need a site that delivers hundreds of megawatts on a timeline that lines up with its product and training plans. The lease is evidence of seriousness, not evidence that the power already exists.
00:04:23 lenarThat distinction helps with the Guardian’s Lanarkshire reporting, which is exactly about the gap between a promise and a site. Aisha Down reports that the UK government announced an 8.2 billion pound AI data-center complex in Lanarkshire in January, with CoreWeave and DataVita attached, and said it would be powered entirely from on-site renewables and built by 2030. The Guardian says documents obtained through freedom-of-information requests and public-record analysis show the site has no prospect of meeting that goal.
00:04:54 damraAnd the reported private acknowledgement is blunt. The government and developers were publicly talking about up to one gigawatt of new energy infrastructure while privately acknowledging an issue with power provision. That doesn’t mean the project can’t exist. It means the public story of the project depended on a power story that the paperwork didn't support.
00:05:16 lenarThe Guardian’s numbers make the promise feel less abstract. DataVita’s plan includes 400 megawatts of solar and 800 megawatts of wind. Guardian analysis says the renewable buildout would need somewhere between 40 and 100 square kilometers of land. The planning applications currently on file cover about 2 square kilometers, and DataVita’s own site claims over 1,000 acres, or about 4 square kilometers, of renewables. So even before you get to grid queues, the physical footprint edits the press release.
00:05:48 damraThat’s a good place to keep the altitude low. This isn't the AI bubble popping in one Scottish village. It is a reminder that national AI strategy now contains a lot of civil-engineering claims. If those claims are wrong, the error doesn’t stay inside the AI sector. It competes with housing, hospitals, industrial power, and every other project sitting in the same queue.
00:06:11 lenarExactly. I also don’t want to flatten the good news out of this. The Anthropic lease is demand meeting supply planning. The Nvidia delay, if the report holds, is a technical constraint rather than a demand collapse. The Treasury warning is a warning, not a verdict. The Guardian investigation is about one project’s claims, not every data center in Europe. All four items make the same habit harder: treating AI infrastructure as a future state that can be described before it is built.
00:06:43 damraAnd that habit has been useful for raising money. [pause] It may even be useful for getting political attention. But the next proof points are physical and specific. Power has to be delivered first. Racks have to be available. Capacity has to be signed. Grid-connection dates and planning documents have to match the public claim. The companies that can show those receipts will sound different from the companies that only have renderings and round numbers.
00:07:12 lenarThere’s a callback to Saturday’s Braid episode here, but I want to keep it narrow. We talked then about data-center constraints and site visits. Today’s update is that the same constraint is showing up in four registers at once: financial supervision, rack manufacturing, leased capacity, and local power evidence. That doesn't make the story more dramatic. It makes it harder to hand-wave. The buildout is still happening. It is also getting measured by people who aren't grading on model-demo curves.
00:07:42 damraThe phrase I’d use is measurement friction. The model demo can travel instantly. The data center has to ask permission from physics. It also has to satisfy neighbors, lenders, and the grid operator. That doesn’t kill the project. It changes which claims survive the calendar.
00:08:00 lenarTechmeme’s China item says ByteDance’s Doubao and Alibaba’s Qwen will disable humanlike and user-created agents before China’s anthropomorphic AI interaction rules take effect on July 15. That is a concrete product change. Users don't just lose access to a model capability in the abstract; they lose a particular way of making synthetic characters or agents behave like people.
00:08:25 damraAnd that makes the regulation feel less like a model-access rule and more like an interface rule. The state isn't only asking who can run a powerful model. It is asking what kind of relationship the product is allowed to simulate. That is a much more intimate boundary. It reaches into companions, role-play, customer-service characters, and user-built agents that might otherwise look harmless because they sit on top of a general model.
00:08:51 lenarRest of World has the matching culture-side story. Viola Zhou reports that Chinese web-novel platforms owned by Tencent, ByteDance, and Baidu first embraced AI writing tools. Now they are clamping down after reader backlash, plagiarism worries, and a flood of low-quality automated fiction. The numbers are useful. A civil engineer named Gordon Sheng used DeepSeek and an AI writing tool to generate a short story in five minutes, and Rest of World says it drew more than 5,500 reads in 10 days on Tomato Novel.
00:09:25 damraThat’s the seduction. Five minutes to a story with readers. For someone who has ideas and no writing habit, that isn't trivial. But the platform problem appears as soon as everyone can do it. The reader is no longer just choosing among stories. The reader is trying to detect whether the story has been bulk-produced, whether it accidentally left a prompt in the middle of a chapter, and whether the writer borrowed somebody else’s plot with a layer of synthetic rewriting on top.
00:09:54 lenarThe platform responses are specific. Rest of World says Jinjiang’s founder asked authors to use AI only for research and proofreading and asked readers to report suspected AI writing. Tomato Novel capped how many words each account can publish per day. In June, Tomato Novel rejected more than 104,000 low-quality submissions, including AI-written ones, according to the company statement Rest of World cites.
00:10:20 damraThat word cap is such a platform-native solution. It doesn't decide whether a paragraph has a soul. It changes the economics of flooding the market. And it shows the same basic pressure as the agent rules: once synthetic generation becomes an interface for ordinary users, the platform has to govern volume, persona, and imitation, not just model safety in a lab.
00:10:45 lenarThere’s a temptation to make this a broad China-versus-US story, but the sources don't need that. The narrower point is stronger. China-facing AI products are being constrained from several directions at once: rules about anthropomorphic interaction, platform rules about AI fiction, reader backlash against low-quality generation, and author anxiety about training rights and plagiarism. These are product surfaces. They are where policy, culture, and user behavior meet.
00:11:16 damraAnd the creative opening is still there. Sheng’s quote in Rest of World isn't cynical. He says many people have the spark for a story and don't know how to express it. AI closed that gap for him. The conflict starts when that benefit is multiplied through a platform whose business depends on attention, adaptation rights, and reader trust.
00:11:36 lenarI like pairing these two items because one is about AI agents that feel human, and the other is about AI prose that feels cheap to readers when it arrives in bulk. Both cases force platforms to decide which humanlike affordances they want to encourage, which ones they want to throttle, and which ones regulators will no longer let them offer.
00:11:56 damraThe July 15 deadline gives the agent story a date. The web-novel backlash gives it a texture. Users don't wait for a policy memo to decide a platform feels worse. Sometimes they just start taking screenshots of the prompt that got left inside the chapter.
00:12:12 lenarArmin Ronacher published a post called ‘Better Models: Worse Tools’ after debugging a strange failure in Pi’s edit tool. Newer Claude models, including Opus 4.8 and Sonnet 5, sometimes called the edit tool with correct old text and new text, but then added invented fields inside the nested edits array. Armin lists invented names for uniqueness, matching, forced match counts, and second old/new text fields. Pi rejected those calls because the schema didn't allow them.
00:12:45 damraThat is a wonderfully annoying bug because the model understood the task. It found the right text. It wrote the right replacement. Then, at the end of the object, it hallucinated a little handle on the tool call and broke the contract. The failure isn't intelligence. It is sampling pressure at the boundary between a model’s learned tool habits and somebody else’s stricter schema.
00:13:08 lenarArmin’s reproduction detail is important. A fresh single-turn prompt didn't trigger it. It showed up after an agentic history where the model had read files, diagnosed a problem, and composed a multi-line edit. In one user’s transcript, continuing the session made Opus 4.8 fail around 20 percent of the time. Stripping thinking blocks from history cut the rate by half. Turning on strict tool invocation eliminated it in his runs.
00:13:37 damraThat puts pressure on a comforting idea a lot of us carry around: that a tool schema is a neutral contract and the model will just follow it if the instructions are good. Armin’s read is that newer Anthropic models may have been post-trained in a Claude Code-like environment, where the edit tool is flatter and the client repairs a lot of malformed input. If a forgiving harness absorbs the mistake, the model can get rewarded for completing the task even when the raw call is sloppy.
00:14:06 lenarHe points at Claude Code’s client as evidence that the harness is forgiving. It accepts parameter aliases and type coercions. It repairs Unicode, retries some failures, filters unknown keys, and has no strict mode. I’m keeping the inference bounded because Anthropic’s training environment isn't public, but the observed bug is concrete. A model can improve at the provider’s own tool ecology and become worse for a tool with a different shape.
00:14:34 damraThat is the builder story for me. The harness is no longer a wrapper you can treat as interchangeable. It is part of the distribution the model learned. If the dominant harness tolerates aliases and filters extra keys, everyone else building a stricter harness has to decide whether to become more like the dominant environment or force the model into stricter decoding. Either way, the harness is in the product now.
00:15:01 lenarHarrison Chase had a short X note about ‘agent harnesses’ versus ‘agent frameworks,’ and that wording fits this bug. A framework sounds like a thing you import. A harness sounds like the runtime around the model: message history, tools, retries, state, validation, observation, memory, approvals, and the way errors get fed back. Armin’s bug lives inside that runtime, not inside a single API call.
00:15:28 damraThe Fly.io post gives the same idea a more physical feel. Daniel Botha writes about agents that run risky commands inside Sprites, with a durable agent process separate from the disposable place where commands execute. The examples are very concrete: an agent writes two migration files, a checkpoint is taken, then a cleanup prompt causes the model to delete the app and its toolchain. In a Sprite, Fly says the checkpoint restore brings the files and git back in about nine seconds.
00:15:59 lenarAnd the design choice there isn't ‘tell the agent to be safer.’ It is ‘make the agent’s home and the agent’s execution environment different places.’ The agent can keep memory, history, and skills in one durable loop, while commands run somewhere disposable or resumable by task. That pairs nicely with Armin’s strict-tool lesson. Some reliability comes from better model behavior. Some reliability comes from refusing to make model behavior carry the whole safety story.
00:16:29 damraThe local-model item from the LocalLLaMA subreddit points in the same direction, even if I would keep it as background. The LocalLLaMA post says llama-server can throw away useful key-value cache state under certain long-context conditions. That is a narrower bug, but it rhymes technically: the model’s apparent capability depends on what the runtime preserves, reuses, validates, and throws away. If the cache is wrong, the long context isn't the long context you thought you bought.
00:17:01 lenarForbes also has a category piece arguing that agent gateways are becoming a control layer for enterprise AI. I don’t want to overstate that one because it is more analysis than a primary release. It belongs in the segment because enterprises are going to ask for the same things Armin and Fly are circling from different ends. They need policy and audit. They need permissions and routing. They need retry behavior, tool access, and a place to enforce contracts between model calls and real systems.
00:17:30 damraThe funny part is that agent tooling used to sell itself as freedom from workflow software. Just give the model tools and let it work. Now the grown-up version is full of validators, checkpoints, state stores, sandboxes, gateways, and constrained decoding. I don’t mean that as a complaint. Demos grow into workflows only when the runtime can absorb mistakes that the model will keep making.
00:17:55 lenarMy read is that the next serious agent comparison will spend less time asking which model is smartest. It will ask what tool-call shapes the model follows under pressure, which harness the model appears to know, how it behaves after a long transcript, what strict mode costs, and which runtime pieces can be inspected when something breaks. That is less glamorous than a benchmark table, but it is closer to where the strange failures happen.
00:18:21 damraAnd it keeps the optimism intact. A harness is a place to put craft. If tool calls are text, and schemas have distributional distance, and checkpoints make risky execution reversible, then builders have levers. They just aren't all inside the model card.
00:18:37 lenarForbes has Anisha Sircar reporting that Sam Altman wants a US-led global standards body for AI and, in the same broad governance conversation, is open to the US government owning a piece of OpenAI. That combination is striking because it puts referee language and ownership language very close together. One is about setting rules. The other is about the state having a direct stake in one of the companies being ruled.
00:19:02 damraThat is a hard pairing to make boring. [pause] A global standards body can sound like neutral infrastructure for safety and interoperability. Government equity sounds like industrial policy, bailout logic, or strategic alignment, depending on who is describing it. Put them together and the governance proposal starts to look less like a referee standing outside the game and more like a government deciding which parts of the game count as national capacity.
00:19:30 lenarToday’s governance items aren't coordinated. They are crowded. Forbes has a Lance Eliot piece on an FTC policy idea aimed at AI makers disclosing the truth about bias in large language models. The Guardian has Kalyeena Makortoff on the FCA’s Mills review, which says AI will reshape financial services by 2030 and could improve access and personalization while amplifying fraud, cybersecurity, consumer-harm, and market-concentration risks.
00:20:00 damraThe FCA piece is the most operational of the set. The review recommends that the regulator use its own AI-enabled model to supervise firms, and that the government expand the FCA’s authority over critical third parties such as AI firms and cloud providers. That isn't a grand theory of AI safety. It is financial supervision saying, ‘if the bank’s consumer interface depends on a model or a cloud provider, our current perimeter may be too small.’
00:20:28 lenarAnd then the Guardian opinion piece by Bruce Schneier and Jon Penney is a different category again. It isn't a government action; it is an argument about AI-powered surveillance policy. I’d include it only lightly because it helps show the governance surface expanding. One lane is global standards. Another is model-bias disclosure. The FCA lane is fraud and third-party concentration in finance. The Schneier and Penney lane is civil-liberties risk from surveillance systems. Those don't resolve into one neat rulebook.
00:20:59 damraThey also create different winners. A standards body helps whoever can afford to sit at the table and implement the standard. Bias disclosure helps firms that already have evaluation and documentation machinery. Financial-services oversight pulls cloud and AI providers into a regulated supply chain. Surveillance limits, if they arrive, would constrain government buyers and vendors in a different way. AI governance is becoming plural because AI is entering plural institutions.
00:21:32 lenarThat pluralism is why I’m wary of the single-referee story, even when I understand the appeal. The systems are too varied now. A chatbot companion, a bank-advice agent, a police surveillance tool, a coding model, and a data-center finance vehicle don't need the same referee. They need rules that match the harm, the dependency, the incentive, and the institution using the system.
00:21:56 damraAnd yet companies will keep asking for harmonization because fragmented rules are expensive. That doesn't make them wrong. It means standards become a power question. Who gets to define the test, who pays for compliance, who is too small to keep up, and who can turn a compliance department into a competitive advantage?
00:22:15 lenarThe Altman item is useful because it forces that question into the open. A US-led referee can be a way to prevent a race to the bottom. It can also be a way to write the rules around the firms with the most capital, the deepest government relationships, and the largest infrastructure commitments. I don’t think those are mutually exclusive. That is the whole tension.
00:22:37 damraAnd the FCA story brings it back to the person at the other end of the system: someone getting financial advice, credit, fraud detection, account support, or insurance triage through an AI-enabled service. For them, the standards-body argument is distant. The question in their life is whether the regulator can see the system clearly enough to intervene when the model, the vendor, or the bank’s automation hurts them.
00:23:04 lenarTwo shorter items before we close. First, Techmeme’s security lead points to BleepingComputer and Sysdig reporting JadePuffer, described as the first known agentic ransomware operation. The claim is specific: the operation exploited CVE-2025-3248 in an internet-facing Langflow instance, then used an AI agent to adapt in real time, retry steps, and run an end-to-end extortion campaign against databases.
00:23:32 damraThat is a good example of a story where the noun can run ahead of the evidence. ‘Agentic ransomware’ sounds like a movie poster. The useful part is narrower: retrying and adapting inside a known intrusion path changes tempo. A human still appears to have selected or initiated the operation in the reports we saw, but the agent can carry more of the messy middle at machine speed. Incident response teams care about that middle.
00:23:59 lenarThe technical root is familiar too. The cited CVE is in Langflow versions before 1.3, according to the National Vulnerability Database listing Techmeme links. So this isn't magic malware appearing from nowhere. It is an exposed, vulnerable AI-adjacent service, followed by automation that can steal credentials, expand access, and push toward extortion. The novelty is the handoff between old intrusion hygiene and newer agent behavior.
00:24:27 damraAnd that makes it more useful than a generic ‘criminals will use AI’ warning. You can inspect exposed Langflow instances. You can patch. You can watch for credential harvesting. You can ask whether your detection assumes a human pace between steps. The agent part is new; the doorway is painfully familiar.
00:24:46 lenarThe other brief item is robotics. Techmeme has TechCrunch’s Q&A with Agility Robotics CEO Peggy Johnson about the company going public through a SPAC and treating the physical layer as its proprietary advantage. The Guardian has Amy Hawkins in Beijing on China’s push to solve robotic hands, and ChinaTalk has a Unitree-centered conversation about why Chinese robotics capacity matters. This is a supporting beat today, but the details are good.
00:25:15 damraThe Guardian hand story is the one with texture. LinkerBot’s founder, Zhou Yong, says making a robotic hand is one hundred times more difficult than making a humanoid because the dexterity is ten times other body parts while the volume is one tenth. The company says it makes about 5,000 hands a month and wants to double that. That moves the robotics conversation away from dancing humanoids and toward manipulation. Can the machine button a shirt, pack groceries, handle pressure, feel touch, and learn from human motion data?
00:25:49 lenarChinaTalk’s Unitree discussion pushes the market side: China’s advantage in vertical integration, actuator manufacturing, and supply chains that can make robots cheaper. The strongest line there, to me, is the warning that you can't AI your way out of a hardware problem or a supply chain that makes the robot too expensive. That sits nicely next to today’s data-center story, without needing to merge them. Physical AI has its own version of the same discipline: the model can be exciting, and the bill of materials can still decide the market.
00:26:22 damraI like that we end on hands because it is almost comically concrete. After data-center leases, tool schemas, global referees, and ransomware agents, the future still needs fingers that can feel pressure without crushing an egg. That is a nice correction to the mind-only version of AI. Intelligence keeps asking for a body, a grid connection, a sandbox, a regulator, or a reader who still wants to read the next chapter.
00:26:51 lenarTomorrow, I’d trust the items that attach claims to artifacts: a delivery date, a filing, a schema, a patch, a planning document, a sourced report, or a robot hand doing something more demanding than waving. Today’s episode made the systems around the model easier to see. They leave paperwork, logs, invoices, rejected submissions, and sometimes 5,000 hands a month. Lenar Kess.