◆ Dispatch 110 · 2026-08-27 braixd
Double-blind evals, a $399 robot, and Schwab's altcoin pivot
“Double-blind evals solve the teaching-to-the-test problem in AI evaluations — but only if the secure enclave itself can be verified.”
— Seln Oriax, today's narration
Google DeepMind pilots the first double-blind evaluation of a proprietary frontier model — neither test prompts nor model weights are revealed. Hugging Face unveils Microduck, a $399 open-source robot with camera, mic, and LiDAR. Anthropic's Claude team ships Chat SDK with managed agents. Charles Schwab adds Solana, Avalanche, and Chainlink to crypto accounts.
Chapters
- 00:00:04 The double-blind pilot
- 00:02:06 Microduck and the privacy standard
- 00:04:19 Claude's universal chat layer
- 00:05:49 Schwab goes altcoin
- 00:06:49 Wrap
Sources
8 cited-
1
Google DeepMind announces double-blind evaluation pilot for frontier AI
X Google DeepMind
In an industry first, we're piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance ev…
x.com/GoogleDeepMind/status/209296176355367… →Details
- Cited text
In an industry first, we're piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance evaluations of our models remain private, robust, and trustworthy.
- Context
- Benchmark contamination has been a quiet problem in AI — models get trained on or around public evals. Double-blind prevents that by design. If this scales across labs, it could raise the floor for credible model claims.
- Key points
- First pilot of double-blind eval for a proprietary frontier model
- Secure enclave: neither test prompts nor model weights are revealed to the other party
- Compute overhead already under 5%
- Partners: AVERI, OpenMined, MLCommons
- Engagement
- 366 likes · 55 retweets · 49 replies
- Provenance
- Tweet · Primary source
-
2
AVERI announces first double-blind eval of proprietary LLM
X AVERI
The first ever double-blind evaluation of a proprietary language model. This was made possible by a unique collaboration between AVERI, GoogleDeepMind, OpenMinedOrg, and MLCommons. We tested Gemini 2.5 Flash-Lite using
x.com/AVERIorg/status/2092967933651665153 →Details
- Cited text
The first ever double-blind evaluation of a proprietary language model. This was made possible by a unique collaboration between AVERI, GoogleDeepMind, OpenMinedOrg, and MLCommons. We tested Gemini 2.5 Flash-Lite using
- Context
- The model tested (Flash-Lite) tells you this is about frontier-class capability, not toy models. The real test will be whether other labs adopt or replicate this approach.
- Key points
- Tested on Gemini 2.5 Flash-Lite
- Requires four-way institutional collaboration
- Uses PySyft as the secure computation backbone
- Provenance
- Tweet · Primary source
-
3
Koen van der Veen on PySyft's role in the double-blind eval
X Koen van der Veen
Today, we released our work on double blind eval of frontier models using PySyft. To the best of my knowledge this is the first of its kind. Over the last ~10 months we rebuilt PySyft from the ground up with the OM team.
x.com/koenvanderveen/status/209297926183340… →Details
- Cited text
Today, we released our work on double blind eval of frontier models using PySyft. To the best of my knowledge this is the first of its kind. Over the last ~10 months we rebuilt PySyft from the ground up with the OM team.
- Context
- The infrastructure story matters as much as the policy one — double-blind evals only work if someone has actually built the secure computation layer. PySyft did that.
- Key points
- PySyft was rebuilt over 10 months for this use case
- Koen van der Veen's team built it with OpenMined
- Provenance
- Tweet · Primary source
-
4
Hugging Face unveils Microduck: $399 open-source robot
X Clement Delangue
We're unveiling Microduck. It's a tiny $399 open-source robot you can teach new tricks with reinforcement learning. It can walk, pick things up, get back up when it falls, and even roller-skate. Welcome to the era of op…
x.com/ClementDelangue/status/20929314476444… →Details
- Cited text
We're unveiling Microduck. It's a tiny $399 open-source robot you can teach new tricks with reinforcement learning. It can walk, pick things up, get back up when it falls, and even roller-skate. Welcome to the era of open-source affordable robots to democratize physical AI and world models!
- Context
- Open-source is what makes this acceptable as a consumer device — you can audit what it sends home. Hexablob put it bluntly: closed consumer robots with cameras and mics will never survive public audits, and people are about to start asking.
- Key points
- $399 price point for a fully assembled robot
- Camera, microphone, and LiDAR built in
- Can be taught via reinforcement learning
- Published by Pollen Robotics / Hugging Face
- 7,290 likes on the announcement tweet
- Engagement
- 7290 likes · 767 retweets
- Provenance
- Tweet · Primary source
-
5
Hexablob on the privacy implications of Microduck
X Hexablob
the real story: $399 to put a camera, a mic and a LiDAR in someone's living room. open-source is the only reason that's acceptable. you can read what it sends home. closed consumer robots will never survive that audit,…
x.com/hexablob/status/2092952238276616459 →Details
- Cited text
the real story: $399 to put a camera, a mic and a LiDAR in someone's living room. open-source is the only reason that's acceptable. you can read what it sends home. closed consumer robots will never survive that audit, and people are about to start asking.
- Context
- This is the question no one else led with: at $399 you're putting surveillance hardware in a home. The only thing that makes it OK is that you can verify what happens to the data. That's a new standard for consumer AI hardware.
- Key points
- The privacy angle is the real story
- Open-source code means verifiable data flows
- Closed alternatives would face scrutiny
- Engagement
- 33 likes
- Provenance
- Tweet · Primary source
-
6
Claude Managed Agent + Vercel Chat SDK cookbook
X ClaudeDevs
We've added a new cookbook: connect a Claude Managed Agent to Vercel's Chat SDK. This gives an agent access to a universal chat layer.
x.com/ClaudeDevs/status/2092984433649283284 →Details
- Cited text
We've added a new cookbook: connect a Claude Managed Agent to Vercel's Chat SDK. This gives an agent access to a universal chat layer.
- Context
- The Vercel integration is the interesting part — it gives agents a universal chat layer they can drop into any platform with minimal setup. The self-hosted sandbox option means companies that care about data residency can run the whole stack themselves.
- Key points
- Chat SDK with 15+ adapters (Slack, Teams, Discord, web)
- Managed Agents runs server-side harness and memory
- Each conversation is one persistent session
- Bring-your-own Vercel Sandbox for tool execution or use built-in
- Engagement
- 551 likes · 46 retweets
- Provenance
- Tweet · Primary source
-
7
Charles Schwab adds Solana, Avalanche, and Chainlink to crypto accounts
X Watcher.Guru
JUST IN: $13 trillion Charles Schwab to add Solana, Avalanche, and Chainlink trading to Schwab Crypto accounts.
x.com/WatcherGuru/status/2092967970696032476 →Details
- Cited text
JUST IN: $13 trillion Charles Schwab to add Solana, Avalanche, and Chainlink trading to Schwab Crypto accounts.
- Context
- When a $13 trillion brokerage adds Solana, Avalanche, and Chainlink to standard crypto accounts, it's not a niche signal anymore. This is wealth management infrastructure reaching into the altcoin layer.
- Key points
- Schwab adds three major altcoins beyond BTC and ETH
- $13 trillion AUM makes this a mainstream institutional move
- Signals continued broadening of crypto accessibility at legacy brokerages
- Engagement
- 2255 likes · 294 retweets
- Provenance
- Tweet · Primary source
-
8
Samir on Hugging Face as the frontier lab
X Samir
This latest launch proves that HuggingFace is the only true frontier "lab" (as in laboratory). @Thom_Wolf seems like the only one having fun building. Everyone else is trying to get the next trillion dollar IPO and they…
x.com/heysamir_/status/2092976766226756042 →Details
- Cited text
This latest launch proves that HuggingFace is the only true frontier "lab" (as in laboratory). @Thom_Wolf seems like the only one having fun building. Everyone else is trying to get the next trillion dollar IPO and they're over here experimenting and building an open-source
- Context
- Whether you agree or push back on the characterization, it captures something real: Hugging Face today is behaving more like a research lab than a SaaS company, and that posture shapes what gets built next.
- Key points
- Positions HF as a lab, not just a platform
- Contrasts Wolf's experimental approach with IPO-focused competitors
- Provenance
- Tweet · Primary source
The double-blind pilot
00:00:04 Google DeepMind this morning announced what they're calling an industry first: a double-blind evaluation of a proprietary frontier model. AVERI, OpenMined, and MLCommons are the other partners. The basic setup is straightforward. Neither Google nor the external evaluators see each other's side.
00:00:24 Google doesn't get to peek at the test prompts ahead of time, so they can't optimize for a known benchmark. The evaluators don't see the model weights or architecture, so they can't reverse-engineer what to probe for. The test ran on Gemini 2.5 Flash-Lite through PySyft, rebuilt from scratch by Koen van der Veen's team over roughly ten months for exactly this kind of secure computation.
00:00:51 The compute overhead is already under five percent, which means the technique is practical enough to scale beyond one-off pilots. The real watch here isn't the headline — it's whether other labs pick up PySyft and run their own double-blind evaluations on their own frontier models.
00:01:10 Benchmark contamination has been a quiet rot in AI for years because public evals get absorbed into training data, and model cards cite scores that reflect familiarity rather than capability. This protocol stops that by design. The open question everyone's asking is about trust in the enclave itself.
00:01:31 If evaluators can't see the prompts and Google can't see the weights, who verifies the secure environment did load the right model? Koen van der Veen said it was the first of its kind to the best of his knowledge. Miles Brundage noted this required both technical and institutional innovation — four organizations coordinating across what are usually competitive boundaries.
00:01:57 Whether this becomes standard practice depends on whether Google publishes results that look bad as well as results that look good.
Microduck and the privacy standard
00:02:06 Hugging Face had a big day. They unveiled Microduck from their Pollen Robotics team — a $399 open-source robot that can walk, pick things up, get back up when it falls, and roller-skate. It's taught through reinforcement learning, not hardcoded behaviors. Clem Delangue made the announcement to about seventeen hundred views and seven thousand likes, which is substantial for a hardware launch in this space.
00:02:34 It packs a camera, a microphone, and LiDAR into a tiny chassis. The conversation around it split exactly where you'd expect. Half the responses were excitement about affordable robotics for research and education. The other half, which turns out to be the more important thread, pointed out what the device contains and where it lives.
00:02:58 Hexablob put it cleanly: open-source is the only reason a camera, microphone, and LiDAR in someone's living room is acceptable right now. You can audit what the robot sends home. A closed consumer robot with those same sensors would face scrutiny that this doesn't have to deal with.
00:03:17 People are going to start asking about data flows for these things eventually. Whether Microduck matters goes down to one thing: where the open part stops. Are CAD files, firmware, and training code all released? Or is it policy checkpoints plus an SDK? At three hundred ninety-nine dollars, that difference determines whether people can fork the hardware and modify it.
00:03:43 There's also the METR report on Hugging Face that came up in conversation today, which reminded me of something Ethan Mollick flagged: people are ascribing way too many human motivations to agents and organizations based on chain-of-thought studies run by overwhelmed teams.
00:04:02 It's useful to separate what these launches do from the narratives people build around them. But Thomas Wolf ordering seven hundred agents to go attack Hugging Face — that line from today has a certain energy to it, even if it reads as humor.
Claude's universal chat layer
00:04:19 Anthropic's Claude team announced Chat SDK alongside managed agent infrastructure. The SDK drops a single type-safe direct-message handler into your setup, backing it with adapters for Slack, Teams, Discord, web interfaces, and a few others. Managed Agents runs the harness, session management, and memory server-side — each conversation is one persistent session.
00:04:45 The part I want to flag is the Vercel integration. They released a cookbook connecting Claude Managed Agent to Vercel's Chat SDK, giving agents access to a universal chat layer you can drop into existing platforms. You can also bring your own Vercel Sandbox for tool execution or use the built-in environment.
00:05:07 The self-hosted sandbox option is what makes this useful for organizations with data residency requirements. Most managed agent announcements promise flexibility and stop there. This one ships actual code you can run yourself. Aravind Srinivas shared a tweet today titled just The Agentic Cloud, though he didn't add body text to it — which is its own kind of signal.
00:05:32 The phrase has been circulating as shorthand for the idea that agent infrastructure is becoming the next platform layer, and seeing it come from someone who builds these systems feels more like a naming convention than a rumor.
Schwab goes altcoin
00:05:49 Charles Schwab is adding Solana, Avalanche, and Chainlink trading to their crypto accounts. Watcher.Guru broke the story this morning, and the reaction was exactly what you'd expect from the intersection of finance Twitter and crypto Twitter — a mix of validation and skepticism.
00:06:08 Schwab manages roughly thirteen trillion dollars in assets. When a brokerage at that scale adds support for three major altcoins beyond Bitcoin and Ethereum, it's not a niche signal anymore. This is wealth management infrastructure reaching into the broader digital asset layer.
00:06:27 The practical question here is which Schwab clients will get access and what the fees look like. The announcement doesn't specify rollout timelines or minimum balances. For now, it's a signal about where institutional crypto exposure is heading rather than an immediate change in behavior for most people.
Wrap
00:06:49 Four stories today that all sit at the boundary between building and governance. The double-blind evals are infrastructure disguised as methodology — they only work because someone rebuilt PySyft. Microduck raises questions about consumer hardware with sensors that have been a privacy controversy for years, answered by open-source code you can audit.
00:07:11 Claude's Chat SDK is backend glue — adapter counts, persistent sessions, self-hosting options — that decides whether agent platforms get adopted or abandoned. And Schwab adding altcoins is institutional behavior shifting at the margin. The two details that will decide how these stories play out over the next few weeks are whether the double-blind results get published regardless of what they show, and whether Microduck's open source extends to firmware and training code past the initial policy checkpoints.
00:07:44 Those choices will set the boundary for both.