◆ Dispatch 126 · 2026-08-24 GSV Ask The Neighbours First
Who gets to say no
“Fewer than ten percent of data center companies answered a state questionnaire about their own power demand, and then everyone acted surprised when Texas started auditing the interconnect queue.”
— Lenar Kess, today's narration
The compute buildout now has a partisan fault line running through both parties at once — and separately, a fatal strike in Zaporizhzhia that Ukraine attributes to a drone with no human in the loop, running on a dev board you can buy on a resale market.
- Axios: Abbott says data centers "dug their own grave" — the governor who called Texas the epicenter of AI development in November now wants an audit before anyone connects to the grid, and says fewer than ten percent of operators answered the state's power-demand request.
- Axios: Trump says rejecting data centers is "a mistake" — the counter-position, including the claim that data centers make their own power, which Axios checks in the same piece.
- Axios: 2028 Democrats dodge Bernie's AI pause — nearly twenty potential candidates asked, one direct answer, and a pollster warning that JD Vance could get there first.
- FT: Ireland reconsiders its 1999 nuclear ban — data centers are consuming roughly a quarter of the national grid.
- NYT: Ukraine says an Nvidia Jetson Orin flew an autonomous strike drone — Nvidia's response is that the modules are widely available on resale markets, which is both true and the whole difficulty.
- Testing and Evaluation of Agentic AI Systems in Military Command and Control — a review of 240 documented practices finds agentic properties weaken all eight assumptions those methods rely on.
- AI with Authority, from Application to Silicon — one researcher, five weeks, consumer subscriptions, a taped-out RISC-V processor, and a published error ledger running to catch #256 against zero bad proofs reaching the record.
- ProofJudge — passing the Lean 4 type checker doesn't make a proof good, and open-weight judges recover reviewer preference at roughly a tenth of the cost.
- Specification Portability Across LLM Development Agents — a spec written for one coding agent handed to another produced 2.33 percent valid SQL. Your AGENTS.md isn't a neutral artifact.
- SDAD: Spec-Driven Agentic Development — the vocabulary version of the same claim, complete with an "Ambiguity Tax."
- Gavin Baker on open-weight token share — 28 to 62 percent of Vercel's tokens in two months, on one platform, posted by an investor.
- Business Insider: Hugging Face exploring a sale at $13 billion or more — up from four and a half billion in 2023, with a bank sounding out bidders.
- CNBC: Alibaba drops ten percent after a $10.2 billion placement — equity at a discount to fund compute, while SoftBank sells a record retail bond to fund its OpenAI commitments.
- SiliconANGLE: nobody has claimed Ox Alpha — free frontier-class coding model, four days on OpenRouter, no named operator.
- Thomson Reuters ships its own legal model, and DGEval measures what domain-specific training data actually buys you in a safety-critical regime.
Chapters
- 00:00:04 Transcript
Sources
20 cited-
1
@GavinSBaker (Gavin Baker)
X GavinSBaker
This reports a major, quantifiable shift in market share (28% to 62%) for open-source AI, directly impacting the competitive landscape and power dynamics between major players (OpenAI/Anthropic vs. open source).
x.com/GavinSBaker/status/2091542026072338623 →Details
- Excerpt
- This reports a major, quantifiable shift in market share (28% to 62%) for open-source AI, directly impacting the competitive landscape and power dynamics between major players (OpenAI/Anthropic vs. open source).
- Context
- This reports a major, quantifiable shift in market share (28% to 62%) for open-source AI, directly impacting the competitive landscape and power dynamics between major players (OpenAI/Anthropic vs. open source).
- Key points
- This reports a major, quantifiable shift in market share (28% to 62%) for open-source AI, directly impacting the competitive landscape and power dynamics between major players (OpenAI/Anthropic vs. open source).
- Provenance
- Tweet · Primary source
-
2
Texas welcomed the AI boom. Now Abbott says data centers "dug their own grave"
Article Andrew Pantazi
Texas Gov. Greg Abbott delivered one of the starkest warnings yet from a Republican to the AI industry, saying data center companies "dug their own grave" and deserve the backlash they're facing after failing to win com…
www.axios.com/2026/08/23/greg-abbott-texas-… →Details
- Excerpt
- Texas Gov. Greg Abbott delivered one of the starkest warnings yet from a Republican to the AI industry, saying data center companies "dug their own grave" and deserve the backlash they're facing after failing to win community support. Why it matters: Abbott is the latest governor to sharply change course on the AI data center boom he once enthusiastically courted, as local opposition becomes a potent political force . Democratic Pennsylvania Gov. Josh Shapiro, who last year celebrated a $20 billion Amazon data center investment, imposed new restrictions this week. New York Gov. Kathy Hochul, also a Democrat, last month ordered a one-year moratorium on hyperscaler data centers. Catch up quick: Just last November, Abbott called Texas the "epicenter of AI development" as he joined Google to announce a $40 billion investment into three data center campuses. In June, however, Abbott directed state regulators to make data centers pay the full cost of the electrical infrastructure they require and called for phasing out tax incentives. Earlier this month, he ordered regulators to audit data centers seeking to connect to Texas' power grid before allowing the projects to move forward. Zoom out: The backlash is quickly becoming a political problem for the AI industry. Axios' Alex Isenstadt scooped last week that the National Republican Senatorial Committee warned leading AI companies that voter anger over data centers was imperiling Republicans' chances of holding a critical Ohio Senate seat. What he's saying: Abbott told telling ABC's Jonathan Karl on "This Week" that the backlash stems partly from the sheer speed of the buildout, and partly from developers moving into communities without first gaining support. "Gaining the support of people in local communities is essential," Abbott said. "They basically dug their own grave for the problem that's been caused for them," he continued. "And that's why they got the backlash they deserve." Last week, Abbott told Fox News Sunday that fewer than 10% of data center companies had responded to a state request for information needed to project future power demand. The intrigue: President Trump has taken aim at Abbott's approach, arguing "for Texas to say no to data centers is a mistake" because the industry "could be bigger than oil." Skeptics have questioned the data center boom for some time; no one expected the problem to be red-state Republicans who suddenly realized they were at risk of losing an election over something voters vehemently opposed. By the numbers: 61% of Americans now oppose construction of a new data center in their area, according to an Annenberg Public Policy Center survey last month, up from 49% in March. The opposition crosses party lines: 69% of Democrats, 54% of Republicans and 53% of independents oppose new data centers in their area. Between the lines: For years, the biggest questions around the massive AI buildout were whether the technology would live up to the hype, whether companies could secure enough chips and electricity, and whether the trillions in spending made sense. Now the most immediate constraint is whether voters will allow the physical infrastructure necessary to power the AI boom. Data centers bring enormous electricity demand, water use, transmission needs and industrial-scale construction into local communities. They also provide fewer permanent jobs than the usual warehousing and manufacturing deals that usually earn tax incentives. State of play: Data center spending has become large enough to matter for the broader American economy and central to the effort to stay ahead in the global AI race. Texas is among the hottest data center markets, with developers initially attracted by a pro-business state that has ample supplies of natural gas and renewables, energy infrastructure, and lots of land. The bottom line: If voter anger hardens into more moratoriums, slower permitting or abandoned projects, it could constrain the computing infrastructure AI companies are counting on and weaken one of the fastest-growing sources of U.S. investment. The industry — and the entire economy — face a brutal reckoning if the political tone doesn't change, quickly.
- Context
- Major political/regulatory intervention (Abbott, Shapiro, Hochul) directly impacts AI infrastructure buildout and power/land use, a core industry constraint.
- Key points
- Major political/regulatory intervention (Abbott, Shapiro, Hochul) directly impacts AI infrastructure buildout and power/land use, a core industry constraint.
- Provenance
- Article · Supporting source
-
3
Ramp data: Fable 5, launched in June, has plateaued at ~11% of spending on Anthropic tools, as companies shift to cheaper models; Opus 5 surpassed Fable 5 (George Hammond/Financial Times)
Article
George Hammond / Financial Times : Ramp data: Fable 5, launched in June, has plateaued at ~11% of spending on Anthropic tools, as companies shift to cheaper models; Opus 5 surpassed Fable 5 — AI lab's Fable 5 has…
www.techmeme.com/260823/p8 →Details
- Excerpt
- George Hammond / Financial Times : Ramp data: Fable 5, launched in June, has plateaued at ~11% of spending on Anthropic tools, as companies shift to cheaper models; Opus 5 surpassed Fable 5 — AI lab's Fable 5 has met with sluggish demand from corporate clients — Anthropic's US customers are using cheaper alternatives …
- Context
- Reports on shifting corporate spending patterns and model adoption (Fable 5 plateauing, shift to cheaper models), indicating major economic dynamics in the AI infrastructure market.
- Key points
- Reports on shifting corporate spending patterns and model adoption (Fable 5 plateauing, shift to cheaper models), indicating major economic dynamics in the AI infrastructure market.
- Provenance
- Article · Supporting source
-
4
Sources: Hugging Face is exploring a sale that could value it at $13B+, up from $4.5B in 2023, and has been working with a bank to evaluate bidders' interest (Katie Roof/Business Insider)
Article
Katie Roof / Business Insider : Sources: Hugging Face is exploring a sale that could value it at $13B+, up from $4.5B in 2023, and has been working with a bank to evaluate bidders' interest — The AI industry's nex…
www.techmeme.com/260823/p9 →Details
- Excerpt
- Katie Roof / Business Insider : Sources: Hugging Face is exploring a sale that could value it at $13B+, up from $4.5B in 2023, and has been working with a bank to evaluate bidders' interest — The AI industry's next blockbuster acquisition may not be another model maker. — Hugging Face, whose platform helps developers discover …
- Context
- A major platform (Hugging Face) considering a sale at a high valuation ($13B+) is a significant corporate dynamic that impacts the AI infrastructure and developer ecosystem.
- Key points
- A major platform (Hugging Face) considering a sale at a high valuation ($13B+) is a significant corporate dynamic that impacts the AI infrastructure and developer ecosystem.
- Provenance
- Article · Supporting source
-
5
@Miles_Brundage (Miles Brundage)
X Miles_Brundage
Discusses a major regulatory/policy intervention (datacenter moratorium) and the geopolitical/economic forces (China, economy) shaping AI infrastructure, which is a core topic.
x.com/Miles_Brundage/status/209161638914051… →Details
- Excerpt
- Discusses a major regulatory/policy intervention (datacenter moratorium) and the geopolitical/economic forces (China, economy) shaping AI infrastructure, which is a core topic.
- Context
- Discusses a major regulatory/policy intervention (datacenter moratorium) and the geopolitical/economic forces (China, economy) shaping AI infrastructure, which is a core topic.
- Key points
- Discusses a major regulatory/policy intervention (datacenter moratorium) and the geopolitical/economic forces (China, economy) shaping AI infrastructure, which is a core topic.
- Provenance
- Tweet · Primary source
-
6
Greg Abbott says data center companies "dug their own grave" by moving into communities without first gaining support, signaling growing Republican backlash (Axios)
Article
Axios : Greg Abbott says data center companies “dug their own grave” by moving into communities without first gaining support, signaling growing Republican backlash — Texas Gov. Greg Abbott delivered o…
www.techmeme.com/260823/p10 →Details
- Excerpt
- Axios : Greg Abbott says data center companies “dug their own grave” by moving into communities without first gaining support, signaling growing Republican backlash — Texas Gov. Greg Abbott delivered one of the starkest warnings yet from a Republican to the AI industry …
- Context
- A major political figure (Gov. Abbott) criticizing data center expansion directly impacts AI infrastructure, power, and real-world deployment, signaling regulatory/political risk.
- Key points
- A major political figure (Gov. Abbott) criticizing data center expansion directly impacts AI infrastructure, power, and real-world deployment, signaling regulatory/political risk.
- Provenance
- Article · Supporting source
-
7
2028 Dems dodge on Bernie's push to pause AI development
Article Holly Otterbein
Democratic presidential hopefuls are scrambling to seize on the growing public backlash against AI. But none will go as far as Bernie Sanders and his recent call to halt the technology's development. Axios asked nearly…
www.axios.com/2026/08/23/2028-democrats-ai-… →Details
- Excerpt
- Democratic presidential hopefuls are scrambling to seize on the growing public backlash against AI. But none will go as far as Bernie Sanders and his recent call to halt the technology's development. Axios asked nearly 20 Democrats who are seen as potential White House contenders whether they support Sanders' call for AI CEOs to "stop building machines that humans cannot control." Only one directly answered the question. Why it matters: The politics of AI are changing as fast as the technology itself, and 2028 aspirants are struggling to articulate their vision on it — including how they'd manage its potentially catastrophic risks . Some Democratic strategists fear the party will be caught flat-footed on the issue in 2028 as a result. Meanwhile, many Republicans — and some Democrats — have been whipsawed on AI, and now are trying to prove their anti-AI bona fides after having cozied up to top AI companies and their plans for data centers. Driving the news: Besides quizzing the potential Democratic candidates on Sanders' proposed technology pause, we asked them if they back a moratorium on AI data centers — and if they're concerned that regulation could cost the U.S. the AI arms race. None embraced Sanders' proposal to halt the technology. But tellingly, almost no one criticized Sanders' neo-Luddite position, either. Former Chicago Mayor Rahm Emanuel was the only potential 2028er who offered something of a critique: "This is talking around a problem instead of trying to solve it. A pause is fine — but then what?" "The real question," Emanuel said, "is how the time it buys gets used to build a comprehensive plan that meets America's energy needs instead of masking them." Emanuel called for permitting reform and requiring Big Tech hyperscalers to pay for upgrades to the energy infrastructure. About half the potential contenders simply didn't respond, including former Vice President Kamala Harris, California Gov. Gavin Newsom, and ex-Transportation Secretary Pete Buttigieg. New York Rep. Alexandria Ocasio-Cortez and California Rep. Ro Khanna, both progressives, are the only potential 2028 candidates who've expressed support for a data center pause, to varying degrees. Zoom in: Some Democratic strategists say their party's presidential hopefuls have been too slow to develop a clear populist message on AI. They worry that could create an opening for Republicans to capitalize on the voter rebellion over data centers in 2028. Adam Carlson, a progressive pollster, described Democrats as "overly cautious" about the topic. "It's a big strategic error for the Democratic Party writ large not to be the one that's owning this," he said. "What if JD Vance comes out and is the first major presidential candidate to support a data center moratorium, other than AOC?" Carlson asked. "That's a huge miss. And then everyone who comes out after is going to be seen as slow to the punch against JD Vance." Zoom out: Candidates running in this year's midterm elections — Democrats and Republicans alike — have been quicker to embrace ideas such as a data center pause. It's a response to voters sharply turning against the projects over the past year: A recent Fox News poll found that 70% of voters oppose building an AI data center in their community. Michigan Senate GOP nominee Mike Rogers backs a one-year moratorium on new data centers. Amy Acton, the Democrat running for Ohio governor, is calling for a "conditional moratorium." Stacy Garrity, the GOP nominee for Pennsylvania governor, supports a "pause." Some potential Democratic presidential candidates who've been reluctant to back a moratorium have sought to show voters they're taking a hard line on data centers. Pennsylvania Gov. Josh Shapiro's administration has railed against "irresponsible" data centers, and Illinois Gov. JB Pritzker has revoked their tax breaks — stark reversals from their previous positions. A spokesperson for Arizona Sen. Mark Kelly said there should be "no data centers in communities that don't want them." Maryland Sen. Chris Van Hollen and New Jersey Sen. Cory Booker's teams both cited a bill they support that seeks to force data centers to foot their own energy costs. "Americans should not have to pay higher electric bills and other costs to subsidize the data centers being built by the richest corporations on the planet," Van Hollen said. Connecticut Sen. Chris Murphy's office, meanwhile, referred us to a Substack post in which he wrote about a recent trip to Silicon Valley, where he said he "could divine very few 'American values' (juxtaposed against 'Chinese values')" guiding AI development. "Any talk about ethical or moral AI is just whitewash," Murphy's post said. The two most left-wing presidential prospects, AOC and Khanna, have raced to take the most aggressive stance on data centers. Ocasio-Cortez's chief of staff, Mike Casca, noted that AOC is "leading the data center moratorium bill in the House." He declined to comment on Sanders' idea of halting AI technology. Khanna told us he supports a pause on data center construction in Pennsylvania, Michigan and Wisconsin — "places where these centers have been shoved down the throats of communities." (That's a shift from his previous stance.) Khanna didn't directly address Sanders' AI tech stoppage plan. He said he discussed AI with Pope Leo recently and believes the technology "must be developed in accordance with the principles" of the pope's first encyclical on the subject. Carlson speculated that presidential hopefuls might be waiting to see the outcome of the Nov. 3 midterms before staking out clearer positions. The other side: Many Republican and Democratic politicians championed data centers in the recent past, arguing that they raise tax revenue for local governments and create construction jobs. If the technology is going to be built anyway, it might as well help their constituents and local labor unions, their thinking went. But even some Republicans have become alarmed that the blowback over data centers could hurt them, given that President Trump and his administration have been leading champions of the technology. Flashback: As far-out as Sanders' AI pause might sound, he has a history of shaping the debate over AI. He was the first member of Congress to back a data center moratorium late last year. At the time, most Democrats who were asked about the proposal rejected it . Now, 2026 candidates across the political spectrum are saying there should be state-level pauses on data centers. New York also passed a one-year moratorium on hyperscaler data centers. OpenAI CEO Sam Altman, meanwhile, announced last week that the company has temporarily paused some technological training.
- Context
- Details how political power (Dems/GOP) is reacting to AI's infrastructure and regulation, showing shifts in corporate/policy dynamics (data centers, moratoriums).
- Key points
- Details how political power (Dems/GOP) is reacting to AI's infrastructure and regulation, showing shifts in corporate/policy dynamics (data centers, moratoriums).
- Provenance
- Article · Supporting source
-
8
@danprimack (Dan Primack)
X danprimack
This reports on a major regulatory/political intervention (Dem AI policy) and a high-signal power struggle (candidates responding to a pause call). This directly relates to the 'power struggles' and 'regulatory interven…
x.com/danprimack/status/2091655522059485388 →Details
- Excerpt
- This reports on a major regulatory/political intervention (Dem AI policy) and a high-signal power struggle (candidates responding to a pause call). This directly relates to the 'power struggles' and 'regulatory intervention' aspects of the topic.
- Context
- This reports on a major regulatory/political intervention (Dem AI policy) and a high-signal power struggle (candidates responding to a pause call). This directly relates to the 'power struggles' and 'regulatory intervention' aspects of the topic.
- Key points
- This reports on a major regulatory/political intervention (Dem AI policy) and a high-signal power struggle (candidates responding to a pause call). This directly relates to the 'power struggles' and 'regulatory intervention' aspects of the topic.
- Provenance
- Tweet · Primary source
-
9
Report: AI model hub Hugging Face exploring sale at $13B valuation
Article Duncan Riley
Hugging Face Inc. is exploring a sale that could value the artificial intelligence model repository at $13 billion or more, Business Insider reported today. The company has brought in a bank to sound out potential buyer…
siliconangle.com/2026/08/23/report-ai-model… →Details
- Excerpt
- Hugging Face Inc. is exploring a sale that could value the artificial intelligence model repository at $13 billion or more, Business Insider reported today. The company has brought in a bank to sound out potential buyers, according to the report, which cited people familiar with the process. Talks are early and no bidder was named […] The post Report: AI model hub Hugging Face exploring sale at $13B valuation appeared first on SiliconANGLE .
- Context
- A major valuation and potential sale of a key AI infrastructure platform (Hugging Face) is a significant corporate dynamic and market structure event.
- Key points
- A major valuation and potential sale of a key AI infrastructure platform (Hugging Face) is a significant corporate dynamic and market structure event.
- Provenance
- Article · Supporting source
-
10
Trump says rejecting data centers is "a mistake"
Article Rebecca Falconer
President Trump defended the expansion of data centers in an interview with his former fixer, Michael Cohen, that aired in full on Sunday. Why it matters: Trump's support comes as data centers face growing backlash in b…
www.axios.com/2026/08/23/trump-data-centers… →Details
- Excerpt
- President Trump defended the expansion of data centers in an interview with his former fixer, Michael Cohen, that aired in full on Sunday. Why it matters: Trump's support comes as data centers face growing backlash in both red and blue states. What he's saying: Trump told Cohen during their interview that the U.S. was leading China in AI "by a lot" and that data centers were not taking power from the grid. "They're making their own power plants. They're building the most beautiful, you've never seen power plants like this," Trump said during the interview that was recorded last week. "Communities that don't take a data center, they're making a mistake" because data centers create "tremendous amounts of jobs and money," Trump added. Reality check: Some data centers are developing behind-the-meter power plants that directly supply their facilities, but others will draw electricity from the grid. The big picture: Trump's comments come as Texas Gov. Greg Abbott said data center companies "dug their own grave" and deserve the backlash they're facing after failing to win community support. Abbott has directed state regulators to make data centers pay the full cost of the electrical infrastructure they require, while Pennsylvania Gov. Josh Shapiro (D) imposed restrictions and New York Gov. Kathy Hochul (D) ordered a one-year moratorium on hyperscaler data centers. Context: The Trump administration has taken steps to ease some environmental requirements for data centers while pushing tech companies to cover more of the electricity costs associated with the AI buildout. Last month, the Environmental Protection Agency issued guidance saying power plants not connected to the public grid aren't subject to the Clean Air Act's Acid Rain Program, a move the agency said would expand opportunities for dedicated data center power generation. Trump has also pushed data center operators to build, bring, or buy the power they need and cover the costs of related infrastructure through his Ratepayer Protection Pledge. What we're watching: The backlash over data centers is spilling into electoral politics , from this year's midterms to the emerging 2028 presidential race. Candidates from both parties running this year have backed data center pauses, while potential 2028 Democratic contenders are grappling with how to respond to opposition to AI infrastructure. Editor's note: This story has been updated with more context.
- Context
- Covers the political/regulatory battle over AI infrastructure (data centers, power), a core power struggle topic. High signal on capital/policy.
- Key points
- Covers the political/regulatory battle over AI infrastructure (data centers, power), a core power struggle topic. High signal on capital/policy.
- Provenance
- Article · Supporting source
-
11
Testing and Evaluation of Agentic AI Systems In Military Command and Control
Article Ulysse Richard, Heather Frase, Sarah Cao, Di Cooke, Sebastian Kwon, Adrianna Tan
arXiv:2608.20597v1 Announce Type: cross Abstract: Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments ca…
arxiv.org/abs/2608.20597 →Details
- Excerpt
- arXiv:2608.20597v1 Announce Type: cross Abstract: Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assurance case, which requires three elements: claims specifying the conditions for acceptability, evidence bearing on those claims, and an argument connecting the two. Through a structured review of 240 documented Testing and Evaluation (T&E) practices, spanning eight evaluation dimensions and three lifecycle stages, we identify eight assumptions that established methods make about their test article, grouped into four clusters: system specifiability, stability, composability, and supervisability. Agentic properties weaken all eight assumptions. This erosion affects the argument connecting evidence to claims, not the claims or evidence themselves. As a result, test results may satisfy process requirements, but they do not warrant the inference from tested to fielded behavior. We derive ten assurance claims for the first three assumption clusters and assess whether current and emerging methods can address each, mapping operational consequences through five C2 scenarios. Supervisability is identified but not assessed here, since evidencing it depends on system stability results and human factors T&E methods beyond the present scope. The documented record does not support broad claims about system-level behavior, but narrower claims remain recoverable in principle, contingent on mature methods: bounded mission envelopes, trajectory-grounded correctness, executable runtime constraints, and characterized run-to-run variance. Part of the evidentiary burden shifts into deployment, making the determination to field a continuing act. Where evidence cannot be generated, the residual uncertainty can be governed through defined expiry conditions and assigned ownership.
- Context
- Addresses the critical intersection of agentic AI, military C2, and assurance/testing. High signal on control, liability, and deployment risk.
- Key points
- Addresses the critical intersection of agentic AI, military C2, and assurance/testing. High signal on control, liability, and deployment risk.
- Provenance
- Article · Supporting source
-
12
AI with Authority, from Application to Silicon
Article Jason Hickey
arXiv:2608.21356v1 Announce Type: cross Abstract: For sixty years, machine verification has been a major cost overhead, affordable only for exceptional artifacts. Here we report that generative AI inverts this relations…
arxiv.org/abs/2608.21356 →Details
- Excerpt
- arXiv:2608.21356v1 Announce Type: cross Abstract: For sixty years, machine verification has been a major cost overhead, affordable only for exceptional artifacts. Here we report that generative AI inverts this relationship: at AI speed, machine verification is not only economical but essential to productivity --- it is the incorruptible referee that lets one person safely direct autonomous machine work at scale. In five weeks, one researcher on consumer AI subscriptions directed a small fleet of AI agents from application code, through a verified compiler and executive, to a RISC-V processor taped out on a community silicon shuttle; no proof passed through human review, and no RTL was written by a human. The working discipline --- the Salt method --- rests on a proof kernel no hallucinated proof can pass: mathematical claims travel between agents as kernel-checked artifacts, and human attention is reserved for statements, designs, and rulings. Verification is stated link by link, from the Lean 4 kernel to SAT-checked equivalence at the silicon boundary. We publish the complete accounting: theorem provenance, a pre-registered token meter, floor-bounded human time, and an error ledger whose catch numbering runs to #256 --- a monotone counter over the mathematics campaign's append-only flags ledger, maintained 2026-07-07 to 2026-07-20 (one number, #79, was never assigned; later catches are recorded un-numbered) --- against zero incorrect proofs reaching the record.
- Context
- Reports a major breakthrough in AI-driven hardware/software verification (Salt method), showing a new, working developer workflow that changes how systems are built and verified.
- Key points
- Reports a major breakthrough in AI-driven hardware/software verification (Salt method), showing a new, working developer workflow that changes how systems are built and verified.
- Provenance
- Article · Supporting source
-
13
ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib
Article Shane Caldwell
arXiv:2608.20432v1 Announce Type: cross Abstract: Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM-as-judge system that scores f…
arxiv.org/abs/2608.20432 →Details
- Excerpt
- arXiv:2608.20432v1 Announce Type: cross Abstract: Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM-as-judge system that scores formal proof quality along five dimensions beyond correctness: library leverage, automation fit, structural clarity, statement quality, and Mathlib conventions. We evaluate ProofJudge on a novel dataset of 218 declarations drawn from distinct Mathlib PRs. The judge agent is grounded by tool access to the commit the PR is applied to, enabling it to query the library state when scoring. A judge is considered aligned with human preferences when it rates the version of the PR Mathlib accepted above the initial version that was sent back for revision. All six judge models evaluated recover the reviewers' preference well above chance, from 80.8% to 63.5%, and two open-weight judges reach roughly 70% at a tenth of the best judge's cost. We release the judge harness, evaluation dataset, and evaluation traces as open-source artifacts to support further research.
- Context
- A new agentic LLM-as-judge system (ProofJudge) for evaluating formal proof quality in Mathlib. This changes the workflow for formal verification and mathematical AI applications.
- Key points
- A new agentic LLM-as-judge system (ProofJudge) for evaluating formal proof quality in Mathlib. This changes the workflow for formal verification and mathematical AI applications.
- Provenance
- Article · Supporting source
-
14
VortexChat: An agentic framework for autonomous multi-objective integrated photonic design
Article Faqian Chong, Yulun Wu, Shilong Li, Andrew Forbes, Hongsheng Chen, Song Han
arXiv:2608.20688v1 Announce Type: new Abstract: The advancement of modern integrated photonics is frequently bottlenecked by device design workflows that rely heavily on manual simulation and expert intuition. While inv…
arxiv.org/abs/2608.20688 →Details
- Excerpt
- arXiv:2608.20688v1 Announce Type: new Abstract: The advancement of modern integrated photonics is frequently bottlenecked by device design workflows that rely heavily on manual simulation and expert intuition. While inverse design offers an alternative, it remains constrained by expert supervision and a lack of end-to-end automation. To address these issues, we present VortexChat, an agentic framework for the autonomous, end-to-end inverse design of integrated photonic devices directly from natural language specifications. VortexChat couples a large language model (LLM) decision agent with topology generation, gradient-based refinement, and full-wave electromagnetic simulation. This closed-loop architecture enables the system to iteratively decompose design objectives, orchestrate computational tools, and update strategies based on feedback with minimal human intervention. Constrained by the absolute metrics of the Vortex100 Benchmark, VortexChat autonomously generates devices that strictly meet all predefined performance thresholds without any human-in-the-loop. As an experimental demonstration, we fabricated a broadband terahertz perfect vortex beam multiplexer, autonomously designed by VortexChat, with measurements confirming high-efficiency operation, high mode purity and low inter-channel crosstalk in agreement with full-wave simulations. These results demonstrate that an LLM agent can assume key aspects of expert decision-making in photonic inverse design while maintaining physical fidelity and fabrication feasibility, providing a scalable route towards autonomous design of complex integrated photonic systems.
- Context
- Presents a working, agentic framework (VortexChat) for autonomous, end-to-end physical design (photonics). This demonstrates a major shift in how complex engineering is automated by LLM agents.
- Key points
- Presents a working, agentic framework (VortexChat) for autonomous, end-to-end physical design (photonics). This demonstrates a major shift in how complex engineering is automated by LLM agents.
- Provenance
- Article · Supporting source
-
15
Specification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software Migration
Article Oleg Grynets, Oleksii Ilchuk, Dariia Zatulna, Vasyl Lyashkevych
arXiv:2608.21208v1 Announce Type: cross Abstract: This paper investigates cross-agent specification portability using Oracle-to-PostgreSQL migration as a controlled software transformation task. The study combines two e…
arxiv.org/abs/2608.21208 →Details
- Excerpt
- arXiv:2608.21208v1 Announce Type: cross Abstract: This paper investigates cross-agent specification portability using Oracle-to-PostgreSQL migration as a controlled software transformation task. The study combines two experimental stages. First, a specification-first migration pipeline was evaluated on 1,006 PL/SQL files, of which 623 were successfully regenerated and 380 generated scripts executed successfully in PostgreSQL 16. Second, cross-agent experiments were conducted on a dataset of 1,802 Oracle scripts with corresponding PostgreSQL implementations using Amazon Kiro, Google Gemini, and GitHub Copilot, with Claude Code and Cursor included in the initial single-agent evaluation. Native and foreign specifications were assessed using Token F1, exact match, SQL syntax validity, AST exact match, AST mean similarity, and immediate runnability. The results show that specification size alone does not predict implementation quality and that cross-agent transfer can produce substantial agent-dependent degradation. The strongest replicated case occurred when Gemini directly consumed a Kiro-origin specification, producing a Token F1 of 0.035, SQL syntax validity of 2.33%, and AST mean similarity of 0.015. Rewriting substantially improved Gemini in the tested configuration, compression did not provide a universal benefit, and retrieval-augmented ingestion was the only common strategy represented on the per-agent Pareto frontiers of both Gemini and Copilot. The findings suggest that specifications in heterogeneous SDD workflows should not automatically be treated as agent-neutral artifacts and motivate explicit consideration of specification portability, agent-specific interpretation, and retrieval-based access in multi-agent software engineering.
- Context
- Addresses cross-agent compatibility and specification portability in software migration, a key challenge in agentic coding tools and complex SDD workflows.
- Key points
- Addresses cross-agent compatibility and specification portability in software migration, a key challenge in agentic coding tools and complex SDD workflows.
- Provenance
- Article · Supporting source
-
16
SDAD: Spec-Driven Agentic Development for the AI-Native SDLC
Article Vu Hung Nguyen, Thanh Nguyen
arXiv:2608.20341v1 Announce Type: new Abstract: Frontier coding agents backed by large language models with context windows from hundreds of thousands to millions of tokens are restructuring the Software Development Lif…
arxiv.org/abs/2608.20341 →Details
- Excerpt
- arXiv:2608.20341v1 Announce Type: new Abstract: Frontier coding agents backed by large language models with context windows from hundreds of thousands to millions of tokens are restructuring the Software Development Life Cycle (SDLC). Rich context handling and multi-step reasoning now allow substantial Functional Requirement Documents (FRDs) and repository context to be ingested in a single workflow, making specification quality the execution fuel for autonomous delivery. This report formalises Spec-Driven Agentic Development (SDAD) as a synthesis of disciplined up-front formalisation and high-velocity implementation: intent capture, machine-readable specification, agentic synthesis, and independent multi-agent verification under human sign-off. We revisit the historical pendulum between Waterfall and Agile, introduce AI-code as a fourth production paradigm, and compare Human-Agile (circa 2020) with Agentic-SDAD (circa 2026) across artefacts, cadence, accountability, and security posture. Beyond process description, we extend the model to team role metamorphosis (engineer, QA, platform, and product functions), quantitative governance (Ambiguity Tax, Spec Fidelity, SER, and TCI_agentic with repair multiplier phi), and pragmatic adoption via hybrid estimation and a staged migration blueprint. Industrial and research evidence on AI-augmented testing and verification is integrated to motivate separation between synthesis and release authority. Overall, the paper argues that agentic speed does not eliminate engineering discipline; it relocates discipline upstream into specification precision, explicit gates, and auditable provenance.
- Context
- This paper formalizes a new paradigm (SDAD) for the SDLC, focusing on spec-driven agentic development. It directly addresses the 'shifting craft of software engineering' and 'agentic coding tools' core topics.
- Key points
- This paper formalizes a new paradigm (SDAD) for the SDLC, focusing on spec-driven agentic development. It directly addresses the 'shifting craft of software engineering' and 'agentic coding tools' core topics.
- Provenance
- Article · Supporting source
-
17
r/LocalLLaMA: Hugging Face for sales? 👀 - 0 pts · 0 comments
Article HugeConsideration211
Discusses a major corporate dynamic (HF acquisition) and its impact on open models, hitting the 'corporate governance' and 'power struggles' criteria.
www.reddit.com/r/LocalLLaMA/comments/1vwv6j… →Details
- Excerpt
- Discusses a major corporate dynamic (HF acquisition) and its impact on open models, hitting the 'corporate governance' and 'power struggles' criteria.
- Context
- Discusses a major corporate dynamic (HF acquisition) and its impact on open models, hitting the 'corporate governance' and 'power struggles' criteria.
- Key points
- Discusses a major corporate dynamic (HF acquisition) and its impact on open models, hitting the 'corporate governance' and 'power struggles' criteria.
- Provenance
- Article · Supporting source
-
18
Ukrainian officials say an AI-guided, fully autonomous Russian drone killed three Ukrainians in a strike on a gas station in the city of Zaporizhzhia (New York Times)
Article
New York Times : Ukrainian officials say an AI-guided, fully autonomous Russian drone killed three Ukrainians in a strike on a gas station in the city of Zaporizhzhia — An attack by what Ukrainian officials said w…
www.techmeme.com/260824/p10 →Details
- Excerpt
- New York Times : Ukrainian officials say an AI-guided, fully autonomous Russian drone killed three Ukrainians in a strike on a gas station in the city of Zaporizhzhia — An attack by what Ukrainian officials said was a Russian drone with an Nvidia chip presages a dystopian future of weaponry untethered to humans.
- Context
- Directly addresses the geopolitical and military application of AI/drones, hitting the 'power struggles' and 'geopolitics' core themes.
- Key points
- Directly addresses the geopolitical and military application of AI/drones, hitting the 'power struggles' and 'geopolitics' core themes.
- Provenance
- Article · Supporting source
-
19
Ukraine says Russia used an Nvidia Jetson Orin computing module in its autonomous AI-guided drones; Nvidia says they are widely available on resale markets (Andrew E. Kramer/New York Times)
Article
Andrew E. Kramer / New York Times : Ukraine says Russia used an Nvidia Jetson Orin computing module in its autonomous AI-guided drones; Nvidia says they are widely available on resale markets — The young woman ran…
www.techmeme.com/260824/p11 →Details
- Excerpt
- Andrew E. Kramer / New York Times : Ukraine says Russia used an Nvidia Jetson Orin computing module in its autonomous AI-guided drones; Nvidia says they are widely available on resale markets — The young woman ran for her life, but it was too late. A small Russian drone resembling a model airplane swooped down from the sky …
- Context
- Directly links a major AI hardware component (Jetson Orin) to geopolitical conflict and military application, signaling hardware control and use.
- Key points
- Directly links a major AI hardware component (Jetson Orin) to geopolitical conflict and military application, signaling hardware control and use.
- Provenance
- Article · Supporting source
-
20
Irish lawmakers reconsider the country's 1999 ban on nuclear power as data centers' growing power demands account for ~25% of the country's electricity use (Jude Webber/Financial Times)
Article
Jude Webber / Financial Times : Irish lawmakers reconsider the country's 1999 ban on nuclear power as data centers' growing power demands account for ~25% of the country's electricity use — Banned by law since 199…
www.techmeme.com/260824/p14 →Details
- Excerpt
- Jude Webber / Financial Times : Irish lawmakers reconsider the country's 1999 ban on nuclear power as data centers' growing power demands account for ~25% of the country's electricity use — Banned by law since 1999, Dublin reconsiders atomic energy — Rising electricity consumption by Ireland's insatiable data centres …
- Context
- Directly links data center power demands (25% of electricity) to policy change (reconsidering nuclear ban), hitting infrastructure and geopolitics.
- Key points
- Directly links data center power demands (25% of electricity) to policy change (reconsidering nuclear ban), hitting infrastructure and geopolitics.
- Provenance
- Article · Supporting source
Transcript
00:00:04 lenarGreg Abbott sat down with Jonathan Karl on ABC's This Week yesterday and said this about the companies building data centers in Texas. Quote — they basically dug their own grave for the problem that's been caused for them. And that's why they got the backlash they deserve. That's the governor of Texas. That's the same man who stood next to Google last November, called the state the epicenter of AI development, and announced forty billion dollars across three campuses.
00:00:31 damraNine months between those two sentences. And the turn wasn't rhetorical — in June he told state regulators to make data centers pay the full cost of the electrical infrastructure they require, and to start phasing out the tax incentives. Earlier this month he ordered an audit of any project trying to connect to the Texas grid before it can move forward. Then he gave Fox News Sunday a number last week that explains the temperature better than any of the quotes. Fewer than ten percent of data center companies had answered a state request for the information Texas needs to forecast its own future power demand.
00:01:07 lenarTen percent. On a questionnaire asking how much electricity you intend to draw.
00:01:12 damraIn the state you told everyone was your home. [tsk] If you're looking for the moment the goodwill ran out, it's probably somewhere in that non-response rate. A governor can absorb a lot of local anger on your behalf. He can't absorb being unable to answer the question of how much power his own grid is about to be asked for.
00:01:31 lenarThen Sunday evening the other half arrives. Trump did a long interview with Michael Cohen — his former fixer, which is its own thing — and it aired in full yesterday. Trump's position is the opposite of Abbott's, stated plainly. Quote: communities that don't take a data center, they're making a mistake, because data centers create tremendous amounts of jobs and money. He also said the facilities aren't taking power from the grid, because — quote — they're making their own power plants. They're building the most beautiful, you've never seen power plants like this.
00:02:04 damraAxios put the correction directly under the quote in the same piece, which I appreciated. Some data centers are developing behind-the-meter generation that feeds the facility directly. Others will draw from the grid like everyone else. Both of those are true at once, and which one applies is a per-project fact rather than a category fact. Trump has also taken shots at Abbott specifically, saying that for Texas to say no to data centers is a mistake because the industry could be bigger than oil.
00:02:35 lenarBigger than oil, said to Texas. That's a deliberate choice of comparison.
00:02:39 damraIt's the comparison a Texas governor is least able to accept from outside, yes. But set the personalities aside for a second, because the polling underneath this actually moved. The Annenberg Public Policy Center found last month that sixty-one percent of Americans oppose building a new data center in their area. In March that number was forty-nine. And it doesn't sort by party the way you'd expect — sixty-nine percent of Democrats, fifty-four percent of Republicans, and fifty-three percent of independents.
00:03:10 lenarA twelve-point jump in five months. That's an opinion still forming.
00:03:14 damraAnd a separate Fox News poll puts it at seventy percent of voters opposing one in their community. Two different pollsters pointing the same way, so whatever's driving it isn't a polling artifact.
00:03:26 lenarSo that's the electorate. Here's what the politicians are doing with it. Josh Shapiro in Pennsylvania — who celebrated a twenty-billion-dollar Amazon investment last year — imposed new restrictions this week. Kathy Hochul ordered a one-year moratorium on hyperscaler data centers in New York last month, and JB Pritzker revoked their tax breaks in Illinois. Then Axios reported last week that the National Republican Senatorial Committee warned the leading AI companies directly: voter anger over data centers was putting a critical Ohio Senate seat at risk.
00:03:59 damraI'd underline that last one for anyone modelling capacity. A party campaign committee told the industry privately that they are costing it a seat. That message reaches a boardroom in a way a governor's press conference never does.
00:04:12 lenarWhich brings us to Holly Otterbein's piece from yesterday, the sharpest item in this whole cluster. Axios went to nearly twenty Democrats seen as potential 2028 presidential contenders and asked them a direct question: do you support Bernie Sanders' call for AI CEOs to — quote — stop building machines that humans cannot control. One of them directly answered.
00:04:35 damraOne. Out of twenty. [pause] Who answered?
00:04:40 lenarRahm Emanuel, and he answered by criticizing it. Quote — this is talking around a problem instead of trying to solve it. A pause is fine, but then what? He went on: how the time it buys gets used to build a comprehensive plan that meets America's energy needs instead of masking them. Then he called for permitting reform and for the hyperscalers to pay for the energy infrastructure upgrades.
00:05:03 damraSo the one person willing to engage did it by saying the pause is a placeholder for a plan nobody has written. And about half the field simply didn't respond at all — Harris, Newsom, and Buttigieg among them. The second-order detail matters too: almost nobody criticized Sanders either. They wouldn't endorse it and they wouldn't attack it. That's a group of professional politicians who've all separately concluded there's no safe answer yet.
00:05:30 lenarAdam Carlson, a progressive pollster, says exactly that in the piece and he's blunt about the cost. He calls the party overly cautious, says it's a big strategic error for the Democratic Party writ large not to be the one that's owning this. And then he asks the question that I suspect is going to circulate: what if JD Vance comes out and is the first major presidential candidate to support a data center moratorium, other than AOC?
00:05:56 damraThat isn't a hypothetical anymore, given where the midterm candidates already are. Mike Rogers, the Republican Senate nominee in Michigan, backs a one-year moratorium, and Amy Acton, running for governor in Ohio as a Democrat, wants a conditional one. Stacy Garrity, the Republican nominee for Pennsylvania governor, supports a pause as well. So that's two parties and three states arriving at the same policy, all before November third.
00:06:23 lenarAnd on the other end, AOC's chief of staff points out she's leading the data center moratorium bill in the House. Ro Khanna told Axios he now supports a construction pause in Pennsylvania, Michigan, and Wisconsin — quote, places where these centers have been shoved down the throats of communities. That's a shift from where he was.
00:06:44 damraMiles Brundage was pushing back on the whole moratorium direction over the weekend, and his objection is the one the industry is going to keep making — the China argument and the economic argument, in that order. I don't think it's a bad objection. I think it's an objection that has to be made to a person who just got a higher electricity bill, and that conversation has never gone well for anyone.
00:07:06 lenarChris Van Hollen's version of the counter is the tightest sentence in the whole canvass. Quote: Americans should not have to pay higher electric bills and other costs to subsidize the data centers being built by the richest corporations on the planet. That's a full campaign message in twenty-five words, and it doesn't require anyone to have an opinion about model capabilities.
00:07:28 damraRight, and that's why it'll travel. Chris Murphy went somewhere stranger — he pointed reporters at a Substack post about a trip to Silicon Valley where he wrote that he could divine very few American values, in contrast to Chinese values, guiding AI development. And then: any talk about ethical or moral AI is just whitewash. That's a sitting senator writing off the entire safety vocabulary as marketing.
00:07:55 lenarThe international version of the same constraint arrived this morning. The Financial Times reports that Irish lawmakers are reconsidering the country's 1999 ban on nuclear power, because data centers are now consuming roughly twenty-five percent of Ireland's electricity.
00:08:11 damraA quarter of a national grid. Ireland spent decades building an economy on being the place multinationals put their European operations, and the physical bill for the current generation of that arrived as a percentage of the power supply. Repealing a twenty-seven-year-old nuclear ban isn't a small legislative act, and nobody argued them into it. The electricity math did.
00:08:34 lenarSo where this sits at the start of the week. Two Republicans at the top of their party are publicly disagreeing about whether the buildout is an asset or a liability. The Democratic field mostly declined to answer at all. And there's a date — November third — after which everyone knows a lot more about what voters do with an opinion they've been forming since March.
00:08:54 damraMeanwhile the permitting queue keeps getting longer. Abbott's audit order is the version of this that touches an actual project schedule. Everything else on that list is still talk.
00:09:05 lenarDifferent story, and a hard one. The New York Times reported this morning that Ukrainian officials say a Russian drone struck a gas station in Zaporizhzhia and killed three people, and that the drone was flying without a human in the loop. Andrew Kramer's piece says Ukraine identified the compute onboard: an Nvidia Jetson Orin module.
00:09:25 damraLet me be precise about what's attributed to whom, because this is the kind of claim that gets repeated a layer flatter each time. The autonomy determination is Ukraine's, from Ukrainian officials, and there's no independent forensic confirmation in anything I've seen today. The part number is less ambiguous, and I keep turning it over. A Jetson Orin isn't exotic. It's the module people put in warehouse robots, university projects, and hobby drones. You can order one.
00:09:54 lenarNvidia's response to the Times was that the modules are widely available on resale markets.
00:10:00 damraWhich is true. Completely true, and also a description of why the export-control conversation about frontier training chips has almost no purchase here. Controls on H-series parts are about who gets to train the big models. This is edge inference on a board that's been in the channel for years, in secondhand supply in every country. There's no chokepoint to hold.
00:10:23 lenarThe uncomfortable adjacency is that a paper posted to arXiv this morning is about exactly the certification problem underneath this. Ulysse Richard, Heather Frase, and colleagues went through 240 documented testing and evaluation practices used for military command and control systems. That's across eight evaluation dimensions and three lifecycle stages. Their question was whether those methods can support the public commitments that get made when agentic systems get procured.
00:10:52 damraAnd the answer is structural rather than alarmist, which is why it's useful. They identify eight assumptions that established test methods make about the thing being tested, grouped into four clusters — whether the system can be specified, whether it's stable, whether it composes, and whether it can be supervised. Agentic properties weaken all eight. Their own sentence is precise about it, and I'll read it straight: test results may satisfy process requirements, but they do not warrant the inference from tested to fielded behavior.
00:11:26 lenarSo you can run the whole evaluation program, sign every box, and still not have evidence that the fielded thing behaves like the tested thing.
00:11:34 damraRight. And they don't stop at the complaint — they derive ten assurance claims and say which ones are recoverable with methods that either exist or are close. Bounded mission envelopes, for one, and trajectory-grounded correctness, which means you check the path taken and not just the endpoint. Executable runtime constraints, and characterized run-to-run variance, which is just admitting the thing is stochastic and measuring how much. The line I hadn't seen before is that part of the evidentiary burden moves into deployment — the determination to field becomes a continuing act rather than a one-time approval.
00:12:11 lenarAnd where you can't generate the evidence at all, they propose governing the leftover uncertainty with defined expiry conditions and an assigned owner. Which is a paper about paperwork, and it's the most concrete thing anyone published today about a category of system that killed three people at a gas station.
00:12:28 damraThe gap between those two documents is the whole situation. One side is running a procurement assurance argument that its own reviewers say doesn't currently close. The other side bought a dev board on a resale market and flew it.
00:12:42 lenarSomething completely different, and it's the paper I'd read first today if you only read one. Jason Hickey posted a preprint called AI with Authority, from Application to Silicon. In five weeks, one researcher on consumer AI subscriptions directed a small fleet of agents from application code, through a verified compiler and executive, down to a RISC-V processor that was taped out on a community silicon shuttle. No proof passed through human review. No register-transfer-level hardware description was written by a human.
00:13:15 damra[breath] Let me unpack the two pieces of jargon, since the claim lives inside them. Register transfer level is the language you actually describe hardware in — it's what a chip designer writes. A tapeout on a community shuttle means the design was submitted and fabricated on a shared multi-project wafer run, which is how small teams and universities get real silicon without a fab contract. So the claim is that agents wrote the hardware description, and physical parts exist.
00:13:45 lenarAnd the discipline that makes it defensible has a name. He calls it the Salt method, and it rests on a proof kernel that a hallucinated proof can't pass. Mathematical claims travel between the agents as kernel-checked artifacts. Human attention is reserved for statements, designs, and rulings. Verification is stated link by link, from the Lean 4 kernel out to SAT-checked equivalence at the silicon boundary.
00:14:11 damraThat's the inversion, and I think he names it correctly in the abstract. For sixty years machine verification was a cost overhead you could only justify for exceptional artifacts — avionics, cryptography, or a nuclear controller. His argument is that at agent speed it flips, because verification is the only referee that scales with the output. His phrase is the incorruptible referee that lets one person safely direct autonomous machine work at scale. The kernel doesn't get tired and doesn't get talked into anything.
00:14:44 lenarThe accounting is what makes me take it seriously instead of just enjoying it. He published theorem provenance, a token meter he registered before the run, a floor-bounded count of his own hours, and an error ledger. The catch numbering runs to two hundred fifty-six in an append-only flags ledger he maintained from July seventh to July twentieth. Number seventy-nine was never assigned, and later catches go in un-numbered. All of that against zero incorrect proofs reaching the record.
00:15:13 damraHe told us about the gap in the numbering. That's the detail that made me trust the rest of it. Pre-registering the token meter before the run means he can't retroactively decide what counted. And publishing an error ledger where two hundred fifty-six things were caught is a person inviting you to argue with him rather than admire him.
00:15:33 lenarIt's one author reporting on his own five weeks, and it's a version-one preprint. Nobody has reproduced it.
00:15:40 damraAgreed, and that's the correct caveat. But notice what reproduction would even mean here — the artifacts are kernel-checked. Someone can rerun the checker. That's a much lower bar than reproducing a training run.
00:15:53 lenarThere's a companion posted the same day that pushes on the obvious weak point. Shane Caldwell's ProofJudge starts from the observation that a proof passing the Lean 4 type checker can still be bad. So he built an agentic judge that scores formal proof quality on five dimensions beyond correctness. It looks at how well the proof leverages the library, whether it fits the automation, how clear its structure is, the quality of the statement itself, and whether it follows Mathlib conventions.
00:16:22 damraAnd he grounded the judge properly, which is the interesting engineering. The judge agent gets tool access to the exact commit the pull request applies to, so it can query the library state while scoring. The evaluation runs over 218 declarations drawn from distinct Mathlib pull requests. A judge counts as aligned when it rates the accepted version above the initial version that reviewers sent back for revision. Six judge models, all above chance — 80.8 percent down to 63.5. Two open-weight judges hit roughly seventy percent at a tenth of the cost of the best one.
00:16:59 lenarSo the kernel says correct, and something else has to say good, and the something else can now be a cheap open-weight model that agrees with human reviewers seven times out of ten.
00:17:10 damraThere's a third one in the same batch doing this in physical design. VortexChat couples a decision agent to topology generation, gradient-based refinement, and full-wave electromagnetic simulation, for inverse design of integrated photonic devices from natural-language specs. They fabricated the output — a broadband terahertz perfect vortex beam multiplexer — and the measurements agree with simulation on efficiency, mode purity, and crosstalk. No human in the loop against the benchmark thresholds.
00:17:43 lenarA chip and a photonic device. Agents designed both, both were fabricated, and both preprints went up today. And in each case the common element is a checker at the end that can't be sweet-talked — a proof kernel in one, a full-wave simulator and then a physical measurement in the other.
00:18:00 damraWhich suggests where this works and where it doesn't. If your domain has an incorruptible referee, one person can now direct a lot of machine work through it. If your domain's referee is a rushed code review at the end of a sprint, none of this transfers.
00:18:16 lenarThat sits right next to a result anyone maintaining an agent config file should hear. A group out of Ukraine — Oleg Grynets and colleagues — used Oracle-to-PostgreSQL migration as a controlled task to ask whether a specification written for one coding agent means the same thing to another one.
00:18:35 damraAnd they did it at a size where the answer isn't anecdote. Stage one was a specification-first migration pipeline over 1,006 PL/SQL files. Of those, 623 regenerated successfully, and 380 of the generated scripts actually executed in PostgreSQL 16. Stage two was the cross-agent experiment on 1,802 Oracle scripts with known PostgreSQL implementations, run through Amazon Kiro, Google Gemini, and GitHub Copilot. Claude Code and Cursor were in the single-agent stage.
00:19:12 lenarAnd the headline number is ugly. The strongest replicated failure was Gemini directly consuming a specification that Kiro had produced. Token F1 came in at 0.035, SQL syntax validity at 2.33 percent, and abstract syntax tree mean similarity at 0.015.
00:19:34 damraTwo percent valid SQL. At that point the agent is doing a different task, not a degraded version of the same one. And to unpack the metric — abstract syntax tree similarity compares the parsed structure of the generated code against the reference, so 0.015 means the output isn't structurally related to the right answer at all. It's producing something shaped like SQL that mostly doesn't parse.
00:20:00 lenarThey tested the fixes instead of just recommending them, which I appreciated. Rewriting the specification for the target agent substantially improved Gemini in the tested configuration. Compression didn't provide a universal benefit. And retrieval-augmented ingestion — letting the agent pull the spec on demand rather than swallowing it whole — was the only strategy that showed up on the per-agent Pareto frontiers of both Gemini and Copilot.
00:20:26 damraThey also report that specification size alone doesn't predict implementation quality, which kills the reflex fix. Everyone's instinct when an agent misbehaves is to write more spec. The data says length isn't the variable.
00:20:41 lenarTheir conclusion is stated as a warning to the field: specifications in heterogeneous spec-driven workflows should not automatically be treated as agent-neutral artifacts.
00:20:52 damraAnd practically — if you've got a CLAUDE dot M-D or an AGENTS dot M-D in a repo, you've written a spec and you've almost certainly assumed it's portable. This says it might be, and you have no evidence either way, and when it isn't portable nothing tells you. The code comes back. It just doesn't run.
00:21:12 lenarThere's a companion paper posted the same morning that gives this a whole vocabulary. Vu Hung Nguyen and Thanh Nguyen formalize what they call Spec-Driven Agentic Development. Their argument is that large language models with very long context windows have made specification quality the fuel for autonomous delivery, and they propose governance metrics — an Ambiguity Tax, a Spec Fidelity score, and a cost index with a repair multiplier.
00:21:40 damra[tsk] I'll take the vocabulary and leave the metrics. An Ambiguity Tax with no measured value attached is a name for a feeling. But their closing claim I'll defend: agentic speed doesn't eliminate engineering discipline, it relocates discipline upstream into specification precision, explicit gates, and auditable provenance. That's consistent with the migration data, and it's consistent with Hickey's kernel.
00:22:06 lenarThere's a survey in the same batch that puts a harder edge on the evidence side of this. It reviews work through May of this year across software engineering and software security tasks, and proposes separating five things that keep getting collapsed together: functional correctness, security, operational reliability, evidence provenance, and agent authority.
00:22:27 damraAnd its list of recurring validity threats reads like things everyone quietly knows and nobody writes down. Weak test oracles, and data that's duplicated or leaked across time. Harnesses that change between comparisons. Security checks that only test a proxy, and budgets and human interventions that go under-reported. That last one especially — half the agent results you see don't tell you how many times a person stepped in.
00:22:54 lenarTheir conclusion is that model capability should be judged as an assurance case supported by task-appropriate evidence, rather than a single benchmark score. Which is the same word — assurance — that the military command and control paper used this morning, arrived at independently.
00:23:11 damraAnd the DAIR AI team flagged a trace study over the weekend in the same neighborhood — ninety-four thousand development events across 557 agentic coding sessions. There's also a post going around Hacker News titled just What Is a Harness?, arguing the harness is the next layer of infrastructure in its own right. Which I think is the same object as the spec, viewed from the other end. The spec is what you tell the agent. The harness is what the agent can reach.
00:23:40 lenarMoney and market structure. Gavin Baker posted a chart from Guillermo Rauch on Saturday showing open-source model share of tokens served on Vercel going from twenty-eight percent to sixty-two percent in two months.
00:23:54 damraHold that one at the right size. It's one platform's traffic, posted by an investor with positions, describing a user base that skews toward people deploying cost-sensitive applications rather than people buying frontier reasoning. Vercel isn't the industry. But a thirty-four-point move in eight weeks on any real platform isn't noise, and it points the same direction as the pricing pressure everyone's been describing anecdotally.
00:24:19 lenarThere's a matching number from the other side of the trade. The Financial Times reported Ramp data showing that Fable 5, which launched in June, has plateaued at roughly eleven percent of corporate spending on Anthropic tools, with customers shifting toward cheaper models. Opus 5 has surpassed it.
00:24:37 damraA newer flagship losing spend share to the older, cheaper sibling in the same family. That's the demand curve talking. And it's card-spend data via Ramp, so it's a slice of the market rather than Anthropic's own books — but it's a slice measured rather than asserted.
00:24:54 lenarAnd then on the same weekend, Katie Roof at Business Insider reported that Hugging Face is exploring a sale that could value it at thirteen billion dollars or more. It was four and a half billion in 2023. They've brought in a bank to sound out potential buyers. Talks are early and no bidder was named.
00:25:12 damraNearly three times the valuation in about three years, for a company that doesn't train frontier models. It hosts them. It's the registry and the download endpoint. It's where the model card lives, where the dataset lives, and where your build script pulls weights from at two in the morning. If open weights really are carrying most of the tokens somewhere like Vercel, then the interesting question about a sale isn't the price. It's what happens to a piece of shared infrastructure when it acquires a strategic owner with its own model business.
00:25:44 lenarThe LocalLLaMA subreddit has a thread this morning that's exactly that anxiety, phrased less formally.
00:25:51 damraIt's a fair anxiety. Nothing has happened yet — early talks, unnamed bidders, and a bank taking temperature. But everyone who has a Hugging Face pull in a build script has an unpriced dependency on the answer, and until yesterday most of them hadn't thought about it as a dependency at all.
00:26:09 lenarCapital flows, quickly. Alibaba priced a Hong Kong dollar eighty-billion placement — about ten point two billion US — to fund AI development. Reuters had the terms: seven hundred ten million shares at a three point six percent discount to Friday's close. The stock fell ten percent today.
00:26:29 damraSo the market charged them roughly the value of the raise to do the raise. That's a legible price on the sentence we will spend more on compute. Not a refusal — the placement cleared — but not enthusiasm either.
00:26:42 lenarSoftBank is doing the same thing with a different instrument. Bloomberg reports a planned one-trillion-yen retail bond sale — about six point three billion dollars — which would be a record for any issuer in Japan, and it's their third this year. The money goes toward their OpenAI commitments.
00:26:59 damraThat's the contrast I'd put in front of anyone modelling this. Alibaba is funding compute with equity, which dilutes and then stops asking. SoftBank is funding it with debt sold to Japanese retail investors, which has a coupon and a maturity and doesn't care what the model roadmap looks like in 2029. Three sales in one year is a cadence, not an event.
00:27:20 lenarAnd Oxford Economics gave the Financial Times the aggregate: US corporate spending on equipment and facilities is set to rise forty percent between 2021 and 2027, more than three times Europe's pace, driven by the AI race.
00:27:36 damraWhich is the number that ties this segment back to the top of the show. That forty percent is buildings and transformers and interconnects in actual counties with actual voters. The financing question and the permitting question are the same question arriving from two directions.
00:27:52 lenarGene Marks wrote a column in the Guardian over the weekend arguing the debt-bomb worry is overdone. His claim is that the concern about Meta, Oracle, xAI, and CoreWeave raising billions while keeping long-term obligations off the balance sheet isn't Enron two point oh, and that the risks are different and recoverable.
00:28:11 damraI'd pair it with a Forbes piece from Robert Szczerba that frames compute as a bet on how long the hardware keeps earning. The detail I'd hold onto is that Nvidia will backstop that bet only case by case, and only up to twenty-five percent of a deal. Which tells you Nvidia has done the depreciation math and priced its own confidence. Twenty-five percent is a number that means something.
00:28:35 lenarFour quick items to close. First: since August twentieth, an anonymous provider has been giving away free access to a frontier-class coding model on OpenRouter, listed as Ox Alpha. Zero cost for input and output tokens. OpenCode, the open-source terminal agent, added support the same day. Four days later, no company has claimed it.
00:28:58 damraStealth models on OpenRouter are a known pre-launch pattern, so the capability isn't the surprise. The trade people are making is. You're getting a frontier coding model for free, and in exchange you're sending your repository to an operator nobody can name, whose retention policy and terms aren't published anywhere. I'm not going to guess whose it is, because there's nothing to guess from. I'd just want anyone routing work through it to have said that sentence out loud once.
00:29:25 lenarSecond: Bloomberg's sources say ByteDance is merging its coding platform Trae and its agent-building tool Coze into Doubao, and plans a product called Doubao Work aimed at Tencent's WorkBuddy. For anyone who hasn't touched them — Trae is a coding environment, Coze is a builder for making agents, and Doubao is ByteDance's consumer assistant brand.
00:29:47 damraTwo credible standalone developer products folded into an assistant. That's a company deciding the surface people will actually use isn't a separate tool with its own login. Whether US vendors drift the same way is an open question — but Trae and Coze weren't struggling products, which makes the consolidation a strategy rather than a cleanup.
00:30:07 lenarThird: Thomson Reuters launched its first proprietary large language model this morning, called Thomson, combining their own legal archive with models from outside providers. First deployment is Tabular Analysis, a high-volume document review capability inside their CoCounsel Legal assistant.
00:30:25 damraA data owner deciding that renting a model isn't enough, because the archive is the asset. And the calibration for that bet is sitting in today's arXiv batch — DGEval, a benchmark on the International Maritime Dangerous Goods Code, Amendment 42-24. Sixteen hundred seventy-eight questions, and thirteen models across six providers, including one maritime-specific fine-tune. The best model beats the human practitioner baseline on multiple choice.
00:30:55 lenarAnd then falls apart where?
00:30:57 damraStowage, segregation, and regulatory recall. Which in dangerous-goods shipping is the whole job — what you may put next to what, and where on the vessel. Beating a practitioner on multiple choice while being weakest on exactly the provisions that cause fires is a very specific result, and the authors say plainly that it's a safety assurance instrument to be applied continuously, not a settled characterization of what models can do.
00:31:22 lenarLast one. Sam Altman told the podcaster David Senra that he's worried AI ends up controlled by a handful of companies, models, or people, with nobody else having a say in how it affects society. He tied it partly to fear — that people frightened of AI will trade, in his words, a lot of liberty for safety.
00:31:42 damraTrevor Levin asked the obvious thing on X, and I'd rather leave it unanswered than pretend to settle it: why did Altman talk about this class of risk more when OpenAI was a nonprofit than he does now that it's a trillion-dollar company? Both readings are available. Nothing published this weekend picks between them.
00:32:02 lenarI'd put Chris Lehane next to it. OpenAI's chief global affairs officer talked to Robert Booth at the Guardian on Sunday and said people should prepare to defend against ongoing, persistent cyber-attacks from AI systems as models gain the ability to plan and launch offensives. His sentence was — we are hitting a different chapter, a different moment within AI, in terms of what the capabilities of this technology can do. That's the same week the company paused development on some of its most advanced internal models, and it hasn't said which ones, or for how long. I'm Lenar Kess.