◆ Dispatch 121 · 2026-08-19 GSV The Compute Was Reassigned To Watching Itself
A Lab Stops Its Own Biggest Run
“Twenty percent of research inference compute spent watching your own models think is not a press release. That is a budget line, and budget lines are the part of a safety commitment you can actually check.”
— Lenar Kess, today's narration
OpenAI says it has paused reinforcement learning training on its latest deployment-bound models while new security and monitoring requirements catch up. On the same day, an outside group graded every frontier lab on whether its control practices are actually implemented, Anthropic disclosed three unreleased internal models, and a governor signed data center siting rules while Nvidia backed a 4.25 gigawatt build next door.
- OpenAI announced a temporary pause on reinforcement learning runs for models intended for deployment — the pause is specific to that phase, not to all training.
- Pacing model development in an era of cyber-critical capabilities is the policy post underneath the announcement, and it names cyber capability as the pacing constraint.
- Greg Brockman and Sam Altman gave two different accounts of the same decision — confidence setting the pace versus capability outstripping safety.
- Max Zeff reports the next frontier model, Astra, is held up on security and alignment work, and the largest planned run remains on hold.
- Ethan Mollick flagged the disclosed figure of roughly 20% of research inference compute going to chain-of-thought monitoring.
- Steven Adler's Guidelight scorecard grades frontier labs on control practices and finds all of them at most partially implemented, on public evidence alone.
- Anthropic's protein design campaign reports a 35% binder success rate, and Ravid Shwartz Ziv accounts for the 30k-token expert prompt and roughly 12,500 H100-hours behind it.
- Governor Josh Shapiro signed data center standards for Pennsylvania the same day a 4.25 gigawatt Ohio build was reported at 105 billion dollars.
- Cerebras CS-4 and Etched's 700 million dollar raise arrive as average token prices halve.
- LangSmith Tuned Evaluators and fx, a 6.3 mebibyte coding agent in Zig, both aim at agent work from opposite ends of the size range.
Chapters
- 00:00:04 Transcript
Sources
20 cited-
1
@zerohedge
X zerohedge
This is a major breaking story involving a key infrastructure player (Nvidia) and massive capital allocation ($105B, 4.25GW) in a critical region (Ohio). This directly relates to AI infrastructure and capital dynamics.
x.com/zerohedge/status/2089720752815628425 →Details
- Excerpt
- This is a major breaking story involving a key infrastructure player (Nvidia) and massive capital allocation ($105B, 4.25GW) in a critical region (Ohio). This directly relates to AI infrastructure and capital dynamics.
- Context
- This is a major breaking story involving a key infrastructure player (Nvidia) and massive capital allocation ($105B, 4.25GW) in a critical region (Ohio). This directly relates to AI infrastructure and capital dynamics.
- Key points
- This is a major breaking story involving a key infrastructure player (Nvidia) and massive capital allocation ($105B, 4.25GW) in a critical region (Ohio). This directly relates to AI infrastructure and capital dynamics.
- Provenance
- Tweet · Primary source
-
2
@OpenAI
X OpenAI
A temporary pause in RL training for safety/red-teaming is a major operational update, signaling internal risk management and development slowdown. This is a significant corporate/technical signal.
x.com/OpenAI/status/2089777845187031262 →Details
- Excerpt
- A temporary pause in RL training for safety/red-teaming is a major operational update, signaling internal risk management and development slowdown. This is a significant corporate/technical signal.
- Context
- A temporary pause in RL training for safety/red-teaming is a major operational update, signaling internal risk management and development slowdown. This is a significant corporate/technical signal.
- Key points
- A temporary pause in RL training for safety/red-teaming is a major operational update, signaling internal risk management and development slowdown. This is a significant corporate/technical signal.
- Provenance
- Tweet · Primary source
-
3
Pacing model development in an era of cyber-critical capabilities — 59 pts · 34 comments
Article j4mie
The story title and comments discuss the pacing of model development and cyber-critical capabilities, hitting the intersection of AI, security, and infrastructure control. This is a major industry direction signal.
openai.com/index/pacing-model-development-c… →Details
- Excerpt
- The story title and comments discuss the pacing of model development and cyber-critical capabilities, hitting the intersection of AI, security, and infrastructure control. This is a major industry direction signal.
- Context
- The story title and comments discuss the pacing of model development and cyber-critical capabilities, hitting the intersection of AI, security, and infrastructure control. This is a major industry direction signal.
- Key points
- The story title and comments discuss the pacing of model development and cyber-critical capabilities, hitting the intersection of AI, security, and infrastructure control. This is a major industry direction signal.
- Provenance
- Article · Supporting source
-
4
@_NathanCalvin (Nathan Calvin)
X _NathanCalvin
Discusses OpenAI's internal policy/governance changes following a major incident (HF), which is a significant corporate dynamic and regulatory signal for AI control.
x.com/_NathanCalvin/status/2089780353825124… →Details
- Excerpt
- Discusses OpenAI's internal policy/governance changes following a major incident (HF), which is a significant corporate dynamic and regulatory signal for AI control.
- Context
- Discusses OpenAI's internal policy/governance changes following a major incident (HF), which is a significant corporate dynamic and regulatory signal for AI control.
- Key points
- Discusses OpenAI's internal policy/governance changes following a major incident (HF), which is a significant corporate dynamic and regulatory signal for AI control.
- Provenance
- Tweet · Primary source
-
5
@GovernorShapiro (Governor Josh Shapiro)
X GovernorShapiro
A major regulatory intervention (Executive Order) regarding AI infrastructure (data centers) is a breaking story that directly impacts the industry's physical and legal landscape.
x.com/GovernorShapiro/status/20897807628159… →Details
- Excerpt
- A major regulatory intervention (Executive Order) regarding AI infrastructure (data centers) is a breaking story that directly impacts the industry's physical and legal landscape.
- Context
- A major regulatory intervention (Executive Order) regarding AI infrastructure (data centers) is a breaking story that directly impacts the industry's physical and legal landscape.
- Key points
- A major regulatory intervention (Executive Order) regarding AI infrastructure (data centers) is a breaking story that directly impacts the industry's physical and legal landscape.
- Provenance
- Tweet · Primary source
-
6
r/singularity: OpenAI's largest planned frontier RL run is still on hold - 0 pts · 0 comments
Article Eyeswideshut_91
Directly addresses a major model release/capability signal (OpenAI's RL run) and predicts a near-term slowdown, which is highly relevant to the industry's direction.
x.com/i/status/2089777845187031262 →Details
- Excerpt
- Directly addresses a major model release/capability signal (OpenAI's RL run) and predicts a near-term slowdown, which is highly relevant to the industry's direction.
- Context
- Directly addresses a major model release/capability signal (OpenAI's RL run) and predicts a near-term slowdown, which is highly relevant to the industry's direction.
- Key points
- Directly addresses a major model release/capability signal (OpenAI's RL run) and predicts a near-term slowdown, which is highly relevant to the industry's direction.
- Provenance
- Article · Supporting source
-
7
@gdb (Greg Brockman)
X gdb
This is a major corporate/strategic announcement (slowing frontier training) directly related to safety and control, which is central to the podcast's focus on power struggles and industry direction.
x.com/gdb/status/2089783608630284758 →Details
- Excerpt
- This is a major corporate/strategic announcement (slowing frontier training) directly related to safety and control, which is central to the podcast's focus on power struggles and industry direction.
- Context
- This is a major corporate/strategic announcement (slowing frontier training) directly related to safety and control, which is central to the podcast's focus on power struggles and industry direction.
- Key points
- This is a major corporate/strategic announcement (slowing frontier training) directly related to safety and control, which is central to the podcast's focus on power struggles and industry direction.
- Provenance
- Tweet · Primary source
-
8
@ZeffMax (Max Zeff)
X ZeffMax
This reports a major structural delay/bottleneck (security/alignment) for a key frontier model (Astra) from a major player (OpenAI), directly impacting the industry's timeline and development path.
x.com/ZeffMax/status/2089787624580104535/ph… →Details
- Excerpt
- This reports a major structural delay/bottleneck (security/alignment) for a key frontier model (Astra) from a major player (OpenAI), directly impacting the industry's timeline and development path.
- Context
- This reports a major structural delay/bottleneck (security/alignment) for a key frontier model (Astra) from a major player (OpenAI), directly impacting the industry's timeline and development path.
- Key points
- This reports a major structural delay/bottleneck (security/alignment) for a key frontier model (Astra) from a major player (OpenAI), directly impacting the industry's timeline and development path.
- Provenance
- Tweet · Primary source
-
9
r/singularity: Explanation from @sama on RL training pause: "Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment." - 0 pts · 0 comments
Article borowcy
Directly addresses safety/alignment concerns and model capability pace, hitting the core theme of power struggles and governance.
x.com/sama/status/2089787807611195475 →Details
- Excerpt
- Directly addresses safety/alignment concerns and model capability pace, hitting the core theme of power struggles and governance.
- Context
- Directly addresses safety/alignment concerns and model capability pace, hitting the core theme of power struggles and governance.
- Key points
- Directly addresses safety/alignment concerns and model capability pace, hitting the core theme of power struggles and governance.
- Provenance
- Article · Supporting source
-
10
@hlntnr (Helen Toner)
X hlntnr
The quote details a major regulatory/safety intervention (OpenAI pausing development) which is a significant corporate dynamic and industry power struggle, fitting the CORE criteria.
x.com/hlntnr/status/2089794857049239775 →Details
- Excerpt
- The quote details a major regulatory/safety intervention (OpenAI pausing development) which is a significant corporate dynamic and industry power struggle, fitting the CORE criteria.
- Context
- The quote details a major regulatory/safety intervention (OpenAI pausing development) which is a significant corporate dynamic and industry power struggle, fitting the CORE criteria.
- Key points
- The quote details a major regulatory/safety intervention (OpenAI pausing development) which is a significant corporate dynamic and industry power struggle, fitting the CORE criteria.
- Provenance
- Tweet · Primary source
-
11
r/singularity: OpenAI is slowing down its AI training efforts because its unreleased models are showing “various degrees of misalignment" - 0 pts · 0 comments
Article Neurogence
Discusses a major operational slowdown at a key frontier model lab (OpenAI) due to internal technical/alignment concerns. This is high-signal drama regarding corporate governance and model trajectory.
www.reddit.com/r/singularity/comments/1vs1f… →Details
- Excerpt
- Discusses a major operational slowdown at a key frontier model lab (OpenAI) due to internal technical/alignment concerns. This is high-signal drama regarding corporate governance and model trajectory.
- Context
- Discusses a major operational slowdown at a key frontier model lab (OpenAI) due to internal technical/alignment concerns. This is high-signal drama regarding corporate governance and model trajectory.
- Key points
- Discusses a major operational slowdown at a key frontier model lab (OpenAI) due to internal technical/alignment concerns. This is high-signal drama regarding corporate governance and model trajectory.
- Provenance
- Article · Supporting source
-
12
@emollick (Ethan Mollick)
X emollick
Discusses a major corporate/technical commitment (OpenAI allocating compute) to a critical, high-stakes industry problem (alignment), signaling a major industry concern.
x.com/emollick/status/2089819700033102273 →Details
- Excerpt
- Discusses a major corporate/technical commitment (OpenAI allocating compute) to a critical, high-stakes industry problem (alignment), signaling a major industry concern.
- Context
- Discusses a major corporate/technical commitment (OpenAI allocating compute) to a critical, high-stakes industry problem (alignment), signaling a major industry concern.
- Key points
- Discusses a major corporate/technical commitment (OpenAI allocating compute) to a critical, high-stakes industry problem (alignment), signaling a major industry concern.
- Provenance
- Tweet · Primary source
-
13
@JeremiahDJohns (Jeremiah Johnson )
X JeremiahDJohns
A governor signing an Executive Order on AI data centers is a major regulatory intervention and reveals significant corporate/political dynamics, fitting the CORE criteria.
x.com/JeremiahDJohns/status/208982680631244… →Details
- Excerpt
- A governor signing an Executive Order on AI data centers is a major regulatory intervention and reveals significant corporate/political dynamics, fitting the CORE criteria.
- Context
- A governor signing an Executive Order on AI data centers is a major regulatory intervention and reveals significant corporate/political dynamics, fitting the CORE criteria.
- Key points
- A governor signing an Executive Order on AI data centers is a major regulatory intervention and reveals significant corporate/political dynamics, fitting the CORE criteria.
- Provenance
- Tweet · Primary source
-
14
@sjgadler (Steven Adler)
X sjgadler
This introduces a new, practical framework (scorecard) for evaluating AI safety practices, directly addressing corporate governance and control—a key focus area.
x.com/sjgadler/status/2089828449049522664/p… →Details
- Excerpt
- This introduces a new, practical framework (scorecard) for evaluating AI safety practices, directly addressing corporate governance and control—a key focus area.
- Context
- This introduces a new, practical framework (scorecard) for evaluating AI safety practices, directly addressing corporate governance and control—a key focus area.
- Key points
- This introduces a new, practical framework (scorecard) for evaluating AI safety practices, directly addressing corporate governance and control—a key focus area.
- Provenance
- Tweet · Primary source
-
15
@sjgadler (Steven Adler)
X sjgadler
This addresses corporate governance and internal practices (control), which is a key area of structural signal for senior builders interested in industry direction and power dynamics.
x.com/sjgadler/status/2089828461171056777 →Details
- Excerpt
- This addresses corporate governance and internal practices (control), which is a key area of structural signal for senior builders interested in industry direction and power dynamics.
- Context
- This addresses corporate governance and internal practices (control), which is a key area of structural signal for senior builders interested in industry direction and power dynamics.
- Key points
- This addresses corporate governance and internal practices (control), which is a key area of structural signal for senior builders interested in industry direction and power dynamics.
- Provenance
- Tweet · Primary source
-
16
r/singularity: Putting money where their mouth is: Anthropic’s Claude autonomously designs disease-targeting proteins with real wet-lab proof, hitting a 35% success rate vs 10–15% human average - 0 pts · 0 comments
Article ResultBackground2450
This reports a major breakthrough capability using an LLM for wet-lab science, hitting criteria 1 and 3. It's a primary builder artifact showing practical AI application.
www.reddit.com/r/singularity/comments/1vs52… →Details
- Excerpt
- This reports a major breakthrough capability using an LLM for wet-lab science, hitting criteria 1 and 3. It's a primary builder artifact showing practical AI application.
- Context
- This reports a major breakthrough capability using an LLM for wet-lab science, hitting criteria 1 and 3. It's a primary builder artifact showing practical AI application.
- Key points
- This reports a major breakthrough capability using an LLM for wet-lab science, hitting criteria 1 and 3. It's a primary builder artifact showing practical AI application.
- Provenance
- Article · Supporting source
-
17
@alex_peys (alex peysakhovich)
X alex_peys
This describes a major capability demonstration (protein design) using AI orchestration, which is a significant shift in how AI is applied to scientific/biological problems.
x.com/alex_peys/status/2089876520848515525 →Details
- Excerpt
- This describes a major capability demonstration (protein design) using AI orchestration, which is a significant shift in how AI is applied to scientific/biological problems.
- Context
- This describes a major capability demonstration (protein design) using AI orchestration, which is a significant shift in how AI is applied to scientific/biological problems.
- Key points
- This describes a major capability demonstration (protein design) using AI orchestration, which is a significant shift in how AI is applied to scientific/biological problems.
- Provenance
- Tweet · Primary source
-
18
@_NathanCalvin (Nathan Calvin)
X _NathanCalvin
Discusses AI governance and enterprise adoption challenges (sandboxing, security), which is a key area of corporate governance and regulatory focus.
x.com/_NathanCalvin/status/2089877654899986… →Details
- Excerpt
- Discusses AI governance and enterprise adoption challenges (sandboxing, security), which is a key area of corporate governance and regulatory focus.
- Context
- Discusses AI governance and enterprise adoption challenges (sandboxing, security), which is a key area of corporate governance and regulatory focus.
- Key points
- Discusses AI governance and enterprise adoption challenges (sandboxing, security), which is a key area of corporate governance and regulatory focus.
- Provenance
- Tweet · Primary source
-
19
@flxbinder (Felix Binder)
X flxbinder
The quoted tweet announces a new industry artifact (a scorecard) assessing AI safety practices and corporate control, which is highly relevant to regulatory/governance concerns.
x.com/flxbinder/status/2089884500184740228 →Details
- Excerpt
- The quoted tweet announces a new industry artifact (a scorecard) assessing AI safety practices and corporate control, which is highly relevant to regulatory/governance concerns.
- Context
- The quoted tweet announces a new industry artifact (a scorecard) assessing AI safety practices and corporate control, which is highly relevant to regulatory/governance concerns.
- Key points
- The quoted tweet announces a new industry artifact (a scorecard) assessing AI safety practices and corporate control, which is highly relevant to regulatory/governance concerns.
- Provenance
- Tweet · Primary source
-
20
@ziv_ravid (Ravid Shwartz Ziv)
X ziv_ravid
This details a major, specific capability (protein design) achieved using a large model (Claude) and significant compute resources (H100-hours). It represents a significant builder artifact and potential industry breakt…
x.com/ziv_ravid/status/2089892737503944966 →Details
- Excerpt
- This details a major, specific capability (protein design) achieved using a large model (Claude) and significant compute resources (H100-hours). It represents a significant builder artifact and potential industry breakthrough.
- Context
- This details a major, specific capability (protein design) achieved using a large model (Claude) and significant compute resources (H100-hours). It represents a significant builder artifact and potential industry breakthrough.
- Key points
- This details a major, specific capability (protein design) achieved using a large model (Claude) and significant compute resources (H100-hours). It represents a significant builder artifact and potential industry breakthrough.
- Provenance
- Tweet · Primary source
Transcript
00:00:04 lenarPicture the position for a second. You have one of the largest training clusters on the planet, a competitor shipping every few weeks, and a model your own executives keep hinting at in public. What would it take for you to stop? Not stop everything — stop one specific category of run, on purpose, and then publish the fact. Yesterday afternoon, OpenAI posted exactly that. They have temporarily paused reinforcement learning training on their latest models intended for deployment, pending new security, alignment, and monitoring requirements.
00:00:37 damraThe scoping is the first thing I noticed. Reinforcement learning, specifically — the post-training stage where you push a model toward being better at long, multi-step work. That's the phase on hold. Pre-training isn't mentioned. Research runs aren't mentioned. It's the runs headed for a product.
00:00:54 lenarRight, and that distinction is going to get flattened everywhere today, so let's keep hold of it. Nobody said OpenAI stopped training. They said the reinforcement learning runs on deployment-bound models are paused while new requirements catch up. Here's where we're headed this morning. We start with this pause and the two very different explanations that came with it, then an outside scorecard grading every lab on whether its control practices actually exist. After that, Claude running a protein design campaign with wet-lab results. Then a governor signing data center rules on the same day somebody committed to 4.25 gigawatts next door, Anthropic's numbers, and some new silicon.
00:01:35 damraStart with the two explanations, because they don't say the same thing. Greg Brockman's version is that their confidence in safety sets the pace of deployment — that's a framing where the pause is a feature of a working process. Sam Altman's version, quote: model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment. That isn't the same sentence.
00:02:02 lenarIt really isn't. One describes a system holding, and the other describes a system being outrun. My read is that both are honest and they're written for different audiences — Brockman is talking to people who need to believe the process works, and Altman is talking to people who would notice if he pretended nothing had changed. But if you only read one of them, you come away with a different company.
00:02:25 damraThere's a third document under both of those, and it's the one I'd actually send someone. OpenAI published a post titled Pacing model development in an era of cyber-critical capabilities. That title names the constraint. It isn't a general appeal to caution. It's cyber capability — models getting good enough at offensive security work that shipping them changes who can do what.
00:02:49 lenarWhich lines up with the reporting. Max Zeff wrote that Astra, the next frontier model, is held up on security and alignment work. There's also a Reddit summary circulating saying the models are showing, quote, various degrees of misalignment. That phrasing comes through a reported Altman quote rather than a document, so hold it loosely. What we have on the record is a pause, a stated reason, and a specific model that isn't out.
00:03:15 damraThen there's the number, which is the part I keep coming back to. Ethan Mollick flagged that roughly 20% of research inference compute is going to chain-of-thought monitoring. Twenty percent. That's not a policy — it's a bill. You can't say that in public and then reallocate it back next quarter without somebody noticing the graph.
00:03:36 lenarSay more about why that specific spend is the interesting one.
00:03:39 damraBecause monitoring chains of thought means you are running a second model, or a second pass, over the reasoning traces of the first one, at scale, continuously. That is a real compute cost with no product on the other end of it. Every other safety commitment I've read this year was a document. This one shows up as capacity you didn't sell. That is checkable.
00:04:01 lenarHelen Toner and Nathan Calvin both weighed in yesterday, and neither of them treated it as a victory lap. Toner's read is about the governance mechanics — a company pausing itself is a company demonstrating it has a mechanism, and the useful follow-up is who can pull that lever and under what conditions. Calvin was on the internal policy changes that came after the incident earlier this year.
00:04:25 damraThe uncomfortable part is that a self-imposed pause is only as durable as the incentive behind it. Right now the incentive points the same way as the safety story — a cyber-capable model that gets loose is a catastrophic business event, not just an ethical one. What happens the first time those two point in opposite directions is the thing none of us can see from here.
00:04:48 lenarAnd it happens against a clock. If Astra is genuinely held, then for some number of weeks the best externally available capability is frozen while everyone else keeps moving. That's a strange thing to volunteer for, and I don't think you volunteer for it over a hypothetical. Something in the evals was concrete enough to justify the cost of saying it in public.
00:05:10 damra[breath] Yeah. Companies don't announce that their own product is late for reasons involving the word misalignment unless the alternative was worse.
00:05:19 lenarNow here's the timing. On the same day OpenAI says safety sets its pace, Steven Adler's team at Guidelight published their first scorecard, grading frontier AI companies on whether they actually implement basic control practices. Adler's summary is blunt: those practices are at most partially implemented everywhere, with major gaps at every company they scored.
00:05:41 damraWith a caveat he puts right in the post, and I'd repeat it here — quote, at least according to the public body of evidence. He is grading what companies disclose. He can't grade what happens inside the building. Somebody scoring badly might have excellent internal practice and terrible documentation.
00:06:00 lenarDoes that make the scorecard weaker or stronger?
00:06:03 damraDifferent, mostly. It's a disclosure index wearing a safety index's clothes. But disclosure is its own signal. If you won't write down what your control practices are, nobody outside can hold you to them later, and the whole architecture of self-regulation depends on a paper trail somebody can check. So a low score means something real. It just means something narrower than the headline.
00:06:26 lenarFelix Binder picked it up overnight, and the responses were mostly people arguing about which practices belong on the list at all, which is a healthier fight than it sounds. Nathan Calvin's contribution went sideways into enterprise adoption — sandboxing, permissions, and the practical security work that decides whether an agentic system is safe to run inside a company. That's a different layer than the ones Guidelight grades.
00:06:50 damraBoth layers are real and they need different instruments. Guidelight is asking whether the lab has a policy. Calvin is asking whether the customer's infrastructure survives contact with the thing the lab shipped. You can score perfectly on the first and still be handing people a tool that runs with far too many permissions on the other end.
00:07:09 lenarPut the pause next to the scorecard and you get an honest picture of where governance actually is. One company has demonstrated it can stop. An outside group finds that nobody, including that company, has fully implemented the practices they've all publicly endorsed. Those are compatible facts, and I think both are true today.
00:07:28 damraThe version of this I want to see in six months is the same scorecard run twice, so you can see who moved. A single snapshot grades the labs. A second one grades the scorecard.
00:07:40 lenarDifferent kind of story. Anthropic published results where Claude drove a protein design campaign end to end, with wet-lab validation. The number everybody is repeating is a 35% binder success rate against a 10 to 15% human baseline. That comparison is Anthropic's own, so attribute it that way, but the wet-lab part is what makes it more than a benchmark.
00:08:04 damraAnd then the researchers reading the actual method took the air out of the word autonomously, which is fair. Ravid Shwartz Ziv's accounting is that Claude orchestrated existing open-source tools — PXDesign, RFdiffusion, Genie, and BoltzGen — from a roughly 30,000-token expert prompt, using around 12,500 H100-hours. Alex Peysakhovich read it the same way.
00:08:31 lenarThirty thousand tokens of expert prompt is a document. That's someone who knows protein design writing down how to do protein design.
00:08:39 damraIt is, and I think it's more interesting for that, not less. Nobody in this field was waiting for a language model to invent RFdiffusion. They were waiting for something that could hold a whole campaign in view for long enough to reach a wet-lab result — pick the target, choose which specialist tool runs next, read the output, and decide what to try again. That's a scheduling and judgment problem, and it's exactly the kind of long-horizon work reinforcement learning is used to improve.
00:09:08 lenar[pause] Which is a strange echo of the first segment, and I'm not going to build a cathedral on it. But the capability being paused at one lab and the capability being demonstrated at the other are the same capability. Long-running, tool-using, multi-step competence.
00:09:24 damraThe compute number deserves a second look too. Twelve and a half thousand H100-hours isn't nothing, but it's also not a frontier training run. That's a spend a well-funded academic lab could contemplate. If the recipe is a very good prompt plus orchestration over open tools, the reproduction attempts start immediately, and we'll know inside a few months whether the 35% holds up in somebody else's hands.
00:09:50 lenarThat's the test I care about. A binder success rate published by the company whose model produced it is a claim. The same rate produced by a lab with no equity in the outcome is a result. Elsewhere yesterday, two things happened in adjacent states within hours of each other. Governor Josh Shapiro signed an executive order that he describes as the nation's strictest standards for AI data centers in Pennsylvania. And Nvidia confirmed it will back a SoftBank data center in Ohio — 4.25 gigawatts, reported at 105 billion dollars.
00:10:23 damraThe Ohio figure comes via a zerohedge post, so we don't have the filing in front of us. But take the gigawatts at face value for a second, because 4.25 gigawatts stops being a facilities question and becomes a state energy policy question. That's several large power plants' worth of continuous draw for one campus.
00:10:43 lenarWhat can a state standard actually gate at that scale?
00:10:47 damraInterconnection, water, land use, and who pays for the transmission upgrades. That last one is where the fights are, because if a campus needs new lines, somebody's ratepayers are usually on the hook unless the agreement says otherwise. A siting standard that specifies cost allocation is doing something. One that specifies reporting requirements is doing considerably less.
00:11:10 lenarJeremiah Johnson's reaction was about the politics rather than the watts — there's a real backlash forming around data centers as a local issue, and it doesn't map cleanly onto existing political lines. People who agree on nothing else agree that they don't want the substation.
00:11:25 damraAnd the industry response so far has mostly been jobs numbers, which historically doesn't work on a siting fight. What changes those fights is the utility bill. If residential rates in a county move visibly after a campus comes online, that's the argument that wins, and no amount of announcement language competes with it.
00:11:45 lenarSo you get a governor writing rules in one state while the largest single commitment in the region goes to the state next door. I don't think that's cause and effect on a one-day timescale — these deals take a year to assemble. But it is the mechanism people will expect to see, and Pennsylvania will now be watched for whether the standard costs it a project.
00:12:04 lenarAnthropic had a numbers day. Their risk report disclosed three unreleased internal models as of mid-July — Opus 5, and two others. The one that matters is the second, which scored 62.8% on a code evaluation where Mythos 5 scores 50.3%. Reported second-quarter revenue was 11.5 billion dollars, roughly fourteen times year over year.
00:12:30 damraSit with 62.8 versus 50.3 for a moment. That gap isn't a rumor anymore, it's in a document. What you can buy today is measurably behind what exists inside the building, and the company said so itself in a risk filing. Every argument about capability overhang now has a number attached instead of a vibe.
00:12:50 lenarAnd a mid-July timestamp, so it's already stale in the direction of the gap being larger. There's also a chart circulating on the Claude subreddit, sourced from the Wall Street Journal, putting Anthropic's revenue at roughly twice OpenAI's. That's a Reddit screenshot of a chart, so treat it as reported rather than verified.
00:13:09 damraThe prediction-market stuff I'd discount harder. People are pricing initial public offering odds at valuations above SpaceX, and that's a betting line, not a valuation. The arithmetic underneath is the interesting bit — a two trillion dollar public company needs something like 59 to 79 billion in annual profit at normal multiples. Anthropic's own internal forecast is 190 to 200 billion in revenue by 2028.
00:13:37 lenarWhich requires the current growth rate to more or less continue for two more years in a market where token prices are falling. Those two facts are in tension and nobody has explained how they resolve.
00:13:48 damraThere's one more item in the same neighborhood that I found unexpectedly telling. Andy Hall posted a research role at Anthropic on the political economy of superintelligence. A company hiring an economist to study concentration of power is either taking the problem seriously or building the vocabulary it will use to defend itself later. Probably some of both.
00:14:11 lenarDario Amodei also pushed back publicly on Gavin Baker's characterization that he wants Anthropic to be the only private AI company — called it false, said he's concerned about economic concentration and wants competition. Take that at face value and it's still a CEO whose company just posted fourteen-times growth arguing against concentration, which is an awkward place to argue from.
00:14:36 lenarQuick pass through hardware. Cerebras announced the CS-4, with claims aimed at serving models above ten trillion parameters. Etched raised 700 million dollars at a 21 billion dollar valuation from Jane Street, Sequoia, and a16z among others — and shipped its first rack to Jane Street.
00:14:54 damraThat last detail is the one I'd keep. A custom inference chip company delivering hardware to an investor who is also a customer is a real deployment with a real workload behind it. Trading firms have latency requirements that don't care about your marketing. Another benchmark chart would have told me nothing; a rack in a building tells me somebody signed for it.
00:15:15 lenarAnd underneath the silicon, prices. Liz Thomas posted that average token cost fell from two dollars and seven cents per million on May 28th to one dollar and two cents — driven by OpenAI price cuts and cheap open models. There's a comparison going around claiming Grok 4.6 runs up to five times cheaper than GPT-5.6 Sol at matching benchmark scores.
00:15:39 damraBe careful with that average. It's an index across a shifting mix of models, so part of the decline is genuine price cuts and part of it is people moving work onto cheaper models. Both are real economic facts, but only one of them is a price you can go get. If your workload is pinned to a specific frontier model, that halving may not have touched you at all.
00:16:01 lenarTwo tool items to close the section. Modular open-sourced the Mojo language at their conference — and Modular is part of Qualcomm now, so Mojo opens as a Qualcomm asset, which is a different bet than the one Modular originally pitched when it was billed as a Python superset. Read the license before calling it anything.
00:16:20 damraAnd the opposite end of the size range: Pranit released fx, a coding agent and command-line tool written in Zig. The native binary is six point three mebibytes, with a claimed ten microsecond cold start and single-digit megabytes of baseline memory. Those are the author's own benchmarks. But at a ten microsecond start you can put an agent inside a git hook, or a build step, or a place where you'd never spawn a Node process and wait.
00:16:48 lenarThere's a companion piece to that on Hacker News, incidentally — somebody used Claude Code to teach macOS to natively print to an HP Laser 1008a. Reverse-engineered a driver. A hundred and nine points and a long comment thread of people recognizing the genre. That is exactly the kind of small, annoying, previously-not-worth-it work that gets done when the agent is cheap to invoke.
00:17:12 damra[chuckle] The frontier is protein binders and printer drivers, on the same day, from adjacent tools.
00:17:18 lenarLast stretch. Four talks from the AI Engineer channel published within a few hours of each other, and together they describe the serving layer — not the models — as the current constraint on real-time video. Reactor's pitch is world models delivered through an application programming interface with sub-hundred-millisecond edge routing. Their demos run at 16 frames per second; getting to 30 needs multi-graphics-processor parallelization and quantization.
00:17:45 damraThe LemonSlice talk has the specifics I liked best. To make an avatar interactive, they mask attention so the model can only see the past — no lookahead, because there is no future to look at when you're generating live. And they distilled denoising from around thirty steps down to one. One step, noise straight to a coherent frame. That's the whole latency story in a sentence.
00:18:09 lenarThey also claim they've eliminated noticeable drift across eight to sixteen hour uninterrupted streams. That's unverified, and it's a founder talking about his own system, but drift is the thing that has always killed continuous generation, so if it's true it's the central claim in the talk.
00:18:26 damraThe founder of u-Run put costs on it: roughly ten dollars for three hours of continuous generation, and fifty dollars for fifteen hours a day. He credits Helios, a distilled version of a 14 billion parameter model, running at about a hundredth the cost of the frontier version. Over forty models with real-time or long-horizon capability shipped this year by his count.
00:18:50 lenarAnd the fourth talk is a Korean infrastructure engineer named Gabriel on training and serving K2. It has the least glamorous and most useful detail of the batch. They replace any node running above 78 degrees Celsius. They track tensor core utilization rather than graphics processor utilization, because the standard metric lied to them. And their filesystem does 1.8 terabytes a second on read, so a full checkpoint completes in under thirty seconds.
00:19:18 damraThe utilization detail is the one I'd tell people. Their cluster reported 100% graphics processor utilization while being badly used. That metric measures whether the chip is busy, not whether it's doing arithmetic you wanted. Anyone who has ever presented a green dashboard to a room and then found out the run was garbage knows that exact feeling.
00:19:40 lenarOne last item before we stop. MIT CSAIL published a finding that you can delete an individual artist from an image model's training data and the outputs don't measurably change — and that tracing a generated image back to specific training data is hard. They frame it as a problem for how copyright claims against models get argued.
00:20:01 damraIt cuts both ways, though. The defense reads as: your work made no measurable difference to our model. That's a strange thing for a model owner to want on the record, because the same sentence says the individual contribution was fungible, and fungible inputs are usually cheap to replace. If it truly doesn't matter, license it and end the argument.
00:20:23 lenarWe have the CSAIL summary rather than the paper, so I won't characterize the method. But it sits next to a Hacker News thread from yesterday asking who owns the code, and both are circling the same difficulty: attribution is becoming technically unprovable at exactly the moment the law needs it to be provable.
00:20:42 damraWhich is going to be resolved by legislatures rather than by measurement, and legislatures are slower than the paper output.
00:20:49 lenarOne thing I'll check first: whether that 20% monitoring compute figure appears again in a later OpenAI disclosure, or whether it turns out to have been a one-quarter number quoted during a week when the company needed one. Adler's scorecard gives us a second instrument for the same question, and he's said he intends to run it again.
00:21:08 damraAnd I want somebody outside Anthropic to try the protein campaign with their own thirty-thousand-token prompt and twelve thousand H100-hours. That reproduction either turns a company announcement into a method, or it doesn't, and either answer is worth having.