◆ Dispatch 042 · 2026-06-16 The Access Tap
The Export Machine Found Its Off Switch
“When a government can close a model overnight, foreign buyers stop treating American AI as a product and start treating it as permission.”
— Jonas Vale, today's narration
Jonas Vale follows the day when U.S. AI export policy collided with export controls, France moved away from Palantir, xAI power became a national-security claim, OpenAI financials surfaced, and agents moved deeper into real-world action.
- Axios on the U.S. AI export strategy and export controls
- Axios on Anthropic, Fable, and cybersecurity leaders
- The Guardian on France, Palantir, and ChapsVision
- TechCrunch on xAI's gas turbines and the DOJ filing
- Ed Zitron on reported OpenAI financials
- OpenAI on deployment simulation research
- The 2026 AI Index report on arXiv
- NVIDIA, CMU, and UC Berkeley's ENPIRE project
- Agent Fair Bench on arXiv
- Bruce Schneier in The Guardian on Anthropic Fable
Chapters
- 00:00:04 The Export Program Found Its Off Switch
- 00:04:10 France Reads The Procurement Fine Print
- 00:08:17 The Data Center Becomes A Security Claim
- 00:12:00 OpenAI Reports Like A Capital Project
- 00:16:16 OpenAI Tests The Crowd Before Release
- 00:20:09 Agents Leave The Screen
Sources
10 cited-
1
Exclusive: OpenAI Losses Increased Nearly 8X in 2025, With Spending Hitting $34 Billion
Article
OpenAI Had $13.07 Billion In Revenue, $34 Billion In Costs and Expenses, and $20.92 Billion In Losses
www.wheresyoured.at/exclusive-openai-financ… →Details
- Cited text
OpenAI Had $13.07 Billion In Revenue, $34 Billion In Costs and Expenses, and $20.92 Billion In Losses
- Key points
- Ed Zitron reports, based on audited financial documents independently verified by the Financial Times, that OpenAI had 13.07 billion dollars in revenue and 34 billion dollars in costs and expenses in 2025.
- The article reports 19.18 billion dollars of R&D spending, 5.73 billion dollars of sales and marketing, and 17.2 billion dollars of expenses to Microsoft.
- Provenance
- Article · Supporting source
-
2
DOJ claims xAI’s unpermitted gas turbines are a matter of national, economic, and energy security
Article
American national, economic, and energy security
techcrunch.com/2026/06/16/doj-claims-xais-u… →Details
- Cited text
American national, economic, and energy security
- Key points
- TechCrunch reports that DOJ sided with xAI in litigation over unpermitted natural gas turbines near Memphis data centers.
- The article cites 57 turbines, NAACP pollution concerns, and SpaceX filing language indicating 2.8 billion dollars of future gas turbine purchases.
- Provenance
- Article · Supporting source
-
3
Trump's fight with Anthropic is now a fight over cybersecurity
Article
They've set a precedent that American models can't do defensive security research.
www.axios.com/2026/06/16/anthropic-fable-tr… →Details
- Cited text
They've set a precedent that American models can't do defensive security research.
- Key points
- Axios reports that Alex Stamos organized a letter signed by nearly 150 security leaders asking the administration to reverse restrictions on Anthropic models.
- The dispute centers on whether proof-of-concept vulnerability work should be treated as dangerous model behavior or defensive security capability.
- Provenance
- Article · Supporting source
-
4
Agentic Robot Policy Self-Improvement in the Real World
Article
reset the scene, execute a policy, verify the outcome, and refine the next iteration
research.nvidia.com/labs/gear/enpire →Details
- Cited text
reset the scene, execute a policy, verify the outcome, and refine the next iteration
- Key points
- NVIDIA, CMU, and UC Berkeley researchers describe ENPIRE, a harness for coding agents to improve robot policies through real-world reset, rollout, verification, and refinement.
- The project reports a 99 percent pass@8 success rate on dexterous manipulation tasks and studies robot utilization, token utilization, and fleet scaling.
- Provenance
- Article · Supporting source
-
5
OpenAI deployment simulation research thread
X
simulating deployment with recent, de-identified user requests
x.com/OpenAI/status/2066969635099144682 →Details
- Cited text
simulating deployment with recent, de-identified user requests
- Key points
- OpenAI described a method for anticipating model behavior before release by simulating deployment on recent de-identified user requests.
- The thread says the method complements red-teaming, uses opted-in ChatGPT data, removes identifiers, and reports aggregate findings.
- Provenance
- Tweet · Primary source
-
6
France to ditch Palantir’s AI data tools in favour of domestic provider
Article
We must use our own AI models
www.theguardian.com/world/2026/jun/16/franc… →Details
- Cited text
We must use our own AI models
- Key points
- France domestic intelligence plans to replace Palantir tools with ChapsVision over several years to avoid dependency on U.S.-controlled tools.
- Prime Minister Sebastien Lecornu announced a 655 million euro AI investment and a state chatbot rollout built on Mistral models.
- Provenance
- Article · Supporting source
-
7
Trump's AI export strategy runs into Trump's export controls
Article
The government's willingness to arbitrarily and abruptly remove America's best models from all foreign use
www.axios.com/2026/06/16/trump-ai-export-st… →Details
- Cited text
The government's willingness to arbitrarily and abruptly remove America's best models from all foreign use
- Key points
- Axios reports that the American AI exports program could be undermined by the same export-control action that forced Anthropic to pull Fable 5 access.
- Applications for the export program are due June 30, 2026, making the unresolved Anthropic dispute a near-term confidence test for U.S. AI export policy.
- Provenance
- Article · Supporting source
-
8
Artificial Intelligence Index Report 2026
Article
Governance frameworks, evaluation methods, education systems, and the data infrastructure needed to track AI impact
arxiv.org/abs/2606.15708 →Details
- Cited text
Governance frameworks, evaluation methods, education systems, and the data infrastructure needed to track AI impact
- Key points
- The 2026 AI Index report says systems for governance, evaluation, education, and impact tracking are struggling to match AI progress.
- The report adds standalone chapters on AI in science and AI in medicine, and includes an analytical framework on AI sovereignty.
- Provenance
- Article · Supporting source
-
9
AgentFairBench: Do LLM Agents Discriminate When They Act?
Article
fairness for LLMs is still measured by grading answers
arxiv.org/abs/2606.16723 →Details
- Cited text
fairness for LLMs is still measured by grading answers
- Key points
- The paper introduces a benchmark for demographic disparity in agent actions across hiring, lending, and medical triage.
- Its pilot used 864 decisions plus replication and found that an arity-mismatched comparison overstated disparity by about 2.4 times.
- Provenance
- Article · Supporting source
-
10
The Anthropic Fable saga proves: we have opened the AI Pandora box. What now?
Article
Fable is just another incremental improvement
www.theguardian.com/commentisfree/2026/jun/… →Details
- Cited text
Fable is just another incremental improvement
- Key points
- Bruce Schneier argues that the Fable issue is not unique to Anthropic and that harnesses can give other models similar capabilities.
- He argues that model bans may only delay the problem and calls for more public, open, and inspectable AI capability and safety choices.
- Provenance
- Article · Supporting source
The Export Program Found Its Off Switch
00:00:04 The Trump administration spent today trying to sell American AI abroad while explaining why one of the best-known American model companies still couldn't sell its newest model abroad. Axios reported this afternoon that the White House's American AI exports program, created by executive order in July 2025, is now colliding with the same export-control move that forced Anthropic to pull Fable 5 access.
00:00:27 The program is supposed to bundle U.S. infrastructure, tools, and models into systems allies can buy. Selected companies get faster export license review, priority access to federal credit programs, government advocacy overseas, and dedicated interagency coordination.
00:00:43 In normal trade-policy language, that is a pretty strong invitation: buy the American stack, and the American state will help you deploy it. Then the American state showed it can interrupt that stack too. Dean Ball, a former Trump administration AI adviser, told Axios that the government's willingness to remove America's best models from all foreign use shows that the export program is no longer relevant to the people making decisions inside the government.
00:01:11 That is a sharp line, and I think it works because it names the buyer's fear. If you're a ministry, bank, hospital system, or telecom in another country, you don't just ask whether the model works. You ask whether access can disappear after a phone call, a political argument, or a classified technical judgment you never get to read.
00:01:30 There is still no public written restoration process for the Anthropic models. That was the follow-up from Monday, June 15, and as of today the answer is still: not yet. Axios says administration officials and Anthropic staff are still trying to resolve the dispute.
00:01:46 The White House defended the action as a balance between innovation and national security, through spokesperson Kush Desai. The Commerce Department's International Trade Administration didn't comment for the Axios story. Applications for the export program are due June 30, so this isn't an abstract policy anxiety with infinite runway.
00:02:06 It is a procurement deadline, and foreign buyers can read a calendar. The cybersecurity side of the fight also got sharper today. In a separate Axios piece, Sam Sabin reported that Alex Stamos organized an open letter signed by nearly 150 security leaders calling on the administration to reverse the restrictions on Anthropic's Fable 5 and Mythos 5.
00:02:27 Stamos told Axios, "They've set a precedent that American models can't do defensive security research." The alleged capability at issue is proof-of-concept vulnerability work, which is exactly the category that makes regulators nervous and defenders useful. If the model can help demonstrate how a bug works, that can help a criminal.
00:02:47 It can also help a maintainer patch the system before the criminal arrives. Katie Moussouris, according to Axios, said the issue didn't involve mass exploitation by the model, but prompts designed to support defensive security work. That distinction matters because technical controls often get worse when they are made under political heat.
00:03:07 You can remove capability from a model. You may also remove the capability that lets a small defender keep up with a better-funded attacker. I don't know which side has the technical record right because we still haven't seen the full record. But I know this much: if the United States wants foreign governments to buy American AI as trusted infrastructure, it has to explain when access can be revoked, who reviews the evidence, and how service comes back.
00:03:34 Bruce Schneier's Guardian essay is useful context because he refuses to treat the ban as a one-company fix. He argues that Fable is an incremental model improvement inside a broader agent harness problem. By harness, he means the ordinary software around the model.
00:03:50 It handles tools, permissions, prompts, routing, and code execution, then turns a model answer into action. If he's right, banning one model buys time at best. Other models can be steered toward similar behavior by better surrounding software, and smaller models can become more capable when they are placed inside a stronger working loop.
France Reads The Procurement Fine Print
00:04:10 France's domestic intelligence service is moving away from Palantir's AI data tools and toward a French provider, ChapsVision. The Guardian reported that Prime Minister Sebastien Lecornu framed the change as a dependency question. His line was blunt: "We must use our own AI models." He also said France couldn't rely on tools developed by foreign powers.
00:04:31 The DGSI contract won't turn over immediately; Palantir's long-term contract was renewed in 2025, so the replacement process may take several years. But the direction is public now, and the reason isn't subtle. Lecornu said France had to avoid depending on partners who could turn off the access tap for artificial intelligence.
00:04:51 That line is hard to separate from the Anthropic fight. Washington's restriction on foreign nationals accessing Anthropic's latest model is now becoming evidence in other countries' procurement arguments. That doesn't mean every European state will suddenly replace U.S.
00:05:07 tools with domestic ones. ChapsVision reported about 200 million euros of revenue in 2025. Palantir reported about 4.5 billion dollars. They aren't in the same industrial weight class. A government can announce autonomy much faster than it can build the software, staffing, compliance, security review, and institutional trust needed to replace an incumbent vendor inside intelligence work.
00:05:31 Still, procurement is where slogans become spending. France isn't only talking about removing one American vendor from one agency. Lecornu announced 655 million euros for AI infrastructure, computing capacity, research, companies, and industrial sectors. France has also started rolling out a government AI tool built on Mistral models to one million of its 2.6 million civil servants.
00:05:54 The stated use cases are practical: speed up legal cases, help researchers secure grants, and reduce the security risk posed by commercial AI tools. There is a dry little irony here. The United States wants to export AI bundles to allies. In the same week it restricts a flagship model, a major ally can point to that restriction and say: this is why our state needs its own stack.
00:06:17 France's argument doesn't have to prove that ChapsVision is better than Palantir tomorrow morning. It only has to persuade French ministers that foreign dependency has a political cost. Once that cost is visible, a domestic vendor can win time, funding, and patience that it might not have received on product quality alone.
00:06:37 Palantir said it would continue to support the French government wherever its solutions are needed. That is the company doing what companies do: preserve the account and wait out the politics. But the broader European pattern isn't just France. The Guardian noted that Germany's BfV internal security service has reportedly selected ChapsVision technology too, Germany's military has said it will no longer use Palantir products, and Britain is reviewing the National Health Service's 330 million pound Palantir data contract after political and parliamentary pressure.
00:07:11 London Mayor Sadiq Khan blocked a proposed 50 million pound Metropolitan Police contract with Palantir on procurement and value-for-money grounds. My read is not that Europe is about to detach from American AI. It is that American AI dependency is now easier to attack in cabinet meetings, parliamentary hearings, and procurement boards.
00:07:31 A tool that used to be evaluated as software now arrives with a question attached: whose state can interrupt this system when the politics change? The dependency argument also has a domestic audience. European governments have been criticized for buying American cloud, American analytics software, and American model access while promising digital sovereignty in speeches.
00:07:54 A switch like this lets a prime minister say the state is converting the speech into procurement. It doesn't settle whether the domestic tool will perform as well. It does make the purchasing officer's risk register look different, because foreign control is no longer a theoretical concern raised by privacy campaigners.
00:08:13 It is now tied to a live American model-access fight.
The Data Center Becomes A Security Claim
00:08:17 The Justice Department sided with xAI in a lawsuit over unpermitted natural gas turbines near the company's Memphis data centers. TechCrunch, citing Wired and the DOJ filing, reported that the department argued the NAACP lawsuit could undermine "American national, economic, and energy security" if it shut off power to AI innovation supporting military operations.
00:08:39 The department also said Grok is one of four AI models supporting mission-critical operations, including recent strikes in Iran. That is a remarkable sentence to see in an environmental dispute over gas turbines. It turns a local permitting and pollution fight into a national-security argument.
00:08:57 The facts on the ground are concrete. The NAACP sued in April after trying for months to stop xAI's use of mobile gas turbines at the Colossus and Colossus 2 data centers. TechCrunch says xAI has added turbines since then, bringing the total to 57. The company argues that because the turbines remain on trailers, they are exempt from Mississippi air pollution regulation for one year.
00:09:21 The Southern Environmental Law Center, filing on behalf of the NAACP, argues that trailer-mounted turbines can still count as stationary under federal law and therefore require regulation. The local claim is health, not vibes. The NAACP says the region was already among the most polluted in the country and has suffered worse air quality since xAI's data centers went online.
00:09:44 TechCrunch lists three pollutants that increased with the turbine count: fine particulate matter, formaldehyde, and nitrogen oxides. Those are associated with asthma and cardiovascular disease, and formaldehyde exposure raises cancer risk. You don't have to settle the legal question to see the distribution problem.
00:10:04 The benefits are national capability, military support, model training, and company valuation. The costs show up in a neighborhood's air. This is the AI infrastructure story without its press-release finish. A data center is land, power, fuel, water, transformers, permits, emissions, and bargaining power.
00:10:22 When the DOJ says a power source supports military operations, it changes the pressure on regulators and courts. A local plaintiff now has to argue not only against a company, but against the national-security value the federal government attaches to that company's compute.
00:10:39 SpaceX's filing adds the forward number. TechCrunch reports that SpaceX said it will buy another 2.8 billion dollars of gas turbines to power AI data centers over the next three years, with at least 2 billion dollars set aside for mobile gas turbines. That isn't a temporary generator borrowed for a weekend.
00:10:58 It is a capital plan. I find this more revealing than another benchmark table. Who gets to convert local air quality into national AI capacity, and under what review? If AI models are now part of military operations, companies will say their power needs deserve deference.
00:11:15 Communities living near the turbines will still breathe the exhaust. The court fight in Memphis is one place where that trade gets priced. The phrase mobile turbine also carries policy weight. Mobile sounds temporary, flexible, almost harmless. The allegation from the environmental side is that a trailer can become stationary in practice when it sits beside a data center and keeps feeding the same load.
00:11:40 That distinction matters because AI companies need power faster than utilities and regulators can usually build or approve it. If temporary generation becomes the bridge, communities are going to ask when the bridge ends, who monitors the emissions, and whether national AI urgency lets companies occupy the exception indefinitely.
OpenAI Reports Like A Capital Project
00:12:00 Ed Zitron published reported OpenAI financial figures today, and the numbers read less like a software company than a capital project with a chat interface. Zitron says he viewed audited financial documents that were independently verified by the Financial Times.
00:12:17 On that basis, he reports that OpenAI had 13.07 billion dollars in revenue in 2025, 34 billion dollars in costs and expenses, and 20.92 billion dollars in operating losses. He also reports a net loss attributable to OpenAI of 38.53 billion dollars, after accounting for conversion-related fair-value changes, noncontrolling interests, and other accounting lines.
00:12:41 That last number needs caution because the structure is complicated. Zitron himself says some parts of the noncontrolling-member treatment are unclear. Still, the expense base is the part that hits you first. The comparison to 2024 is useful. OpenAI reportedly had 3.7 billion dollars in revenue in 2024, 12.48 billion dollars in costs and expenses, and a 5.09 billion dollar net loss attributable to the company.
00:13:08 In 2025, revenue grew sharply, but costs grew into a much larger machine: 7.5 billion dollars in cost of revenue, 19.18 billion dollars in research and development, 5.73 billion dollars in sales and marketing, and 1.57 billion dollars in general and administrative expenses.
00:13:26 At the end of 2025, OpenAI reportedly had just over 50 billion dollars in assets, with almost half in cash. The Microsoft lines are the institutional story. Zitron reports that OpenAI paid Microsoft 10.59 billion dollars for research and development expenses in 2025, likely tied to training costs, plus 6.047 billion dollars related to cost of revenue, 527 million dollars for sales and marketing, and 42 million dollars in general and administrative expenses.
00:13:56 Total expenses to Microsoft came to 17.2 billion dollars. He also reports that SoftBank paid OpenAI 867 million dollars in 2025, while Microsoft paid OpenAI 303 million dollars. That is a strange dependency map. Microsoft is investor, cloud provider, distribution partner, and major recipient of OpenAI spending.
00:14:16 OpenAI is a product company, a research lab, and a buyer of enormous compute. The public experiences the company through ChatGPT and APIs, but the financial record describes a firm converting capital into model capability at breathtaking speed. I wouldn't treat Zitron's interpretation as the final account of OpenAI's business because he is a known critic of the company, and he says more reporting is coming.
00:14:43 But the reported documents matter because audited figures are harder to wave away than vibes about an AI bubble. If the numbers hold, OpenAI has demand. Thirteen billion dollars of revenue is demand. Serving it requires training, inference, distribution, and sales costs at enormous scale.
00:15:02 This also changes how the model-access and export-policy fights read. A frontier lab spending tens of billions has to find customers with budgets large enough and durable enough to support that machine. Governments, banks, cloud partners, defense users, and major enterprises are not side markets.
00:15:21 They are how the math starts to work. So when export controls, procurement sovereignty, and cloud dependency show up in the same week as the financials, they are the business model meeting the institutions that can pay for it. There is also a labor and sales signal inside the expense mix.
00:15:40 Spending nearly 6 billion dollars on sales and marketing in 2025 says OpenAI isn't waiting for the product to sell itself. It is building an enterprise distribution machine while paying for the compute machine. That matters because enterprise AI revenue usually comes with support, security review, procurement cycles, and promises about uptime and data handling.
00:16:04 Those promises are expensive, and they pull the company deeper into the institutions that will later demand explanations when model behavior, pricing, access, or compliance changes.
OpenAI Tests The Crowd Before Release
00:16:16 OpenAI shared research today on simulating model deployment before release, using recent de-identified user requests to study candidate model responses. The company's X post says it analyzed only ChatGPT conversations from users who allow their data to be used to improve models.
00:16:32 OpenAI said it removed account-linked identifiers and identifiable information before analysis, and reported only aggregate findings. It described the method as a complement to traditional evaluations and red-teaming, especially for rare or severe risks. The purpose is to estimate how often unwanted behavior may occur in realistic use and to surface new behavior before a model reaches production.
00:16:56 That sounds technical, but the governance question is straightforward. The best data for predicting public model behavior is often public model use. External evaluators usually don't have that data. OpenAI said as much in the thread: deployment simulation works best with representative production data, which outside evaluators often can't access.
00:17:17 That is a serious asymmetry. A lab can see the distribution of real user requests, the messy long tail, and the failure patterns that show up only after millions of people start asking normal and abnormal questions. A regulator, academic evaluator, or civil-society group usually sees a much narrower slice.
00:17:36 OpenAI also said it extended the method to agentic deployments with stateful tools, where tool simulators can produce realistic trajectories when given enough context and capabilities. That part matters for liability. A chatbot answer can be wrong and embarrassing.
00:17:52 An agent with tools can send an email, query a database, book a service, change a setting, or make a recommendation that an institution acts on. Testing those systems with single-turn prompts is like testing a logistics company by asking whether a driver can identify a truck in a photograph.
00:18:09 It tells you something, but it doesn't tell you how the route behaves under time pressure, bad weather, and a dispatcher who keeps changing the order. The 2026 AI Index report, posted on arXiv today, makes the larger point in institutional language. Its abstract says governance frameworks, evaluation methods, education systems, and the data infrastructure needed to track AI's impact are struggling to match the pace of the technology.
00:18:36 The report also says this year's edition tracks more ambitious testing across reasoning, safety, and real-world task execution, while those measurements are becoming harder to rely on. It adds standalone chapters on AI in science and AI in medicine, which is a useful signal about where the measurement burden is moving.
00:18:55 I don't dislike OpenAI's method. In fact, this is closer to the evidence you would want before a release: take real usage patterns, anonymize them, run candidate models through the likely traffic, and look for changes in behavior. But the same method makes the oversight problem sharper.
00:19:12 If the most informative test requires private production data, then public trust depends on how much of the method, aggregate result, and failure category the company is willing to publish. The phrase "trust us, we tested it on real usage" won't be enough when the model is acting inside schools, clinics, call centers, and government workflows.
00:19:33 This is also where privacy and safety pull against each other. The safest release evidence may come from the most sensitive data the company has: real conversations, real workflows, and real attempted misuse. OpenAI says it used opted-in data and removed identifiers.
00:19:49 Good. But when this becomes the release standard, outsiders will ask how representative the opted-in data is, which categories were excluded, and whether the riskiest users are the least likely to consent. A deployment simulation can be better than a toy benchmark and still miss the people who most change the risk profile.
Agents Leave The Screen
00:20:09 NVIDIA, CMU, and UC Berkeley researchers published a project called ENPIRE, which turns coding agents loose on real robot policy improvement. The claim is unusually concrete. The project describes a loop where the system resets the scene, executes a policy, verifies the outcome, and refines the next iteration.
00:20:27 Coding agents propose algorithmic changes, run trials on physical robots, inspect logs and failures, consult literature, and edit the training and control code. The project reports a 99 percent pass-at-eight success rate on dexterous manipulation tasks such as Push-T, organizing pins into a pin box, cutting a zip tie, and GPU insertion.
00:20:46 Jim Fan summarized the demo as eight Codex agents with a robot fleet, GPUs, and a token budget trying to solve tasks quickly while keeping the robots busy and safe. I would keep the awe and the caution in the same hand here. It is striking to see software agents close the loop through hardware, not just through code tests or simulated environments.
00:21:07 The project site also gives the constraint that makes this different from another agent benchmark: before a robot policy can improve itself, the task has to become self-resetting and self-verifying. Someone has to define how the physical scene gets returned to a known randomized state.
00:21:23 Someone also has to define how the outcome is scored without a human standing there, and how failures are captured in a way the agent can use. That is where a lot of the labor is hiding. The project is also direct about resource pressure. It defines mean robot utilization and mean token utilization, then notes that coding agents don't fully use robot resources while they read logs, write code, debug, or wait for the model.
00:21:48 Larger fleets can reach success sooner, but they also spend more tokens because agents read peer branches, summarize results, and coordinate. In other words, physical scaling isn't free. You can accelerate the search by adding robots and agents, but the bill moves into hardware, GPU time, tokens, lab operations, and safety review.
00:22:08 A separate paper posted today, Agent Fair Bench, asks a less cinematic question: do large language model agents discriminate when they act? The authors point out that fairness for language models is often measured by grading answers, while agents increasingly take actions: screening applicants, recommending credit, and triaging patients.
00:22:27 Their benchmark uses synthetic matched profiles across hiring, lending, and medical triage, changing only a name-coded race-and-gender signal. The pilot ran 864 decisions plus a test-retest replication. Its methodological warning is useful: comparing a six-group score spread against a two-run noise difference overstated disparity by about 2.4 times.
00:22:48 Against a better matched noise floor, Claude Haiku 4.5 showed no demographic effect above sampling noise in that pilot. That result isn't a universal fairness certificate. It is a reminder that once agents act, the audit has to measure actions, thresholds, tool calls, and noise.
00:23:04 A model can sound fair and still route people differently. It can also look biased if the benchmark compares the wrong quantities. The institutional consequence is plain enough: as agents leave the chat window for robots, clinics, hiring screens, and credit workflows, the evaluation has to follow them into action.
00:23:22 One more piece of the ENPIRE project shouldn't get lost in the demo: the system needs automatic verification. That is the difference between an agent that claims progress and a robot lab that can measure it. In software, a test can fail in milliseconds and the rollback is often cheap.
00:23:39 In hardware, a bad attempt can bend a pin, wear an actuator, misplace a tool, or leave the next trial in a bad starting state. Verification is not just an evaluation metric here. It keeps the next experiment from inheriting an unobserved physical mistake. By June 30, written export rules either appear, or foreign buyers learn that model access is a revocable favor.
00:24:00 Jonas