Archive BRAIXD
The local compute bet, persistent agents, and benchmark signals / DISPATCH 113
PDF RSS

Dispatch 113 · 2026-08-31 braixd

The local compute bet, persistent agents, and benchmark signals

/ 00:08:55 / 12 sources

“Apple doesn't need the Mac to hold the most powerful model in the universe. It just needs local models to be good enough for a huge share of the work that people actually do.”

— Seln Oriax, today's narration

Apple rebuilt its entire desktop Mac line around local AI — 512 GB of unified memory at the top tier. At the same time, Google published a paper on persistent agents with skill wikis, François Chollet is pushing back on unverified benchmark claims, DHH opens the Omacom Foundation to corporate patronage, and we look at what went wrong when OpenAI's age-verification system decided a 27-year-old was under 13.

Chapters

  1. 00:00:04 Apple's Mac line, rebuilt around local AI
  2. 00:03:23 Google's WikiSkill and persistent agents
  3. 00:04:55 Chollet on benchmark integrity
  4. 00:06:09 Omacom Foundation goes corporate
  5. 00:07:37 When automation makes you under 13

Sources

12 cited
  1. 1

    OpenAI age verification system deletes a 27-year-old's account, claims they're under 13

    Article Ok_Kaleidoscope_9721 — Reddit user posting from r/OpenAI about their personal account deletion experience.

    A 27-year-old user's OpenAI account was deleted because the system mistook the issuance date on their Taiwanese ID for a birth date, interpreting them as born in 2014. The appeal went through multiple loops without reso…

    www.reddit.com/r/OpenAI/comments/1w3bk4e/im… →
    Details
    Excerpt
    A 27-year-old user's OpenAI account was deleted because the system mistook the issuance date on their Taiwanese ID for a birth date, interpreting them as born in 2014. The appeal went through multiple loops without resolution.
    Context
    Automated age-verification systems are still error-prone enough to permanently destroy user relationships over a single data point. The friction compounds when each resolution step feeds back into the same automated pipeline.
    Key points
    • OpenAI's system read an ID issuance date (Sep 5, 2014) as the user's birth date
    • Account deleted for being 'under 13'; Trust & Safety confirmed the error but the account was later deleted again after automated review
    • User completed adult verification twice without getting confirmation of completion
    • The final appeal decision came in five minutes via email — no reason, no reference number
    Provenance
    Article · Supporting source
  2. 2

    DHH announces Omacom Foundation corporate patronage

    Thread David Heinemeier Hansson (DHH) — DHH is the creator of Ruby on Rails and CTO of 37signals (Basecamp). He has a history of building foundations aligned with his tools — Rails Foundation is a prior example.

    DHH opens the Omacom Foundation to corporate patronage. 1Password and 37signals become Distinguished Corporate Patrons, each committing $100,000/year over three years. Total pot reaches $12.6M. 'Open source doesn't need…

    x.com/dhh/status/2094422653217984617 →
    Details
    Excerpt
    DHH opens the Omacom Foundation to corporate patronage. 1Password and 37signals become Distinguished Corporate Patrons, each committing $100,000/year over three years. Total pot reaches $12.6M. 'Open source doesn't need to be a commune.'
    Context
    Three-year corporate commitments at $100k/year are a different class of commitment than one-time sponsorships or donation buttons. It's infrastructure as sustained business cost, not philanthropy.
    Key points
    • Omacom Foundation is now open for corporate patronage with 1Password and 37signals as founding Distinguished Corporate Patrons
    • Each company commits $100,000/year over three years; the total pot reaches $12.6M
    • DHH: 'Open source doesn't need to be a commune' — signaling an explicit break from grassroots-only funding norms
    • The fund supports Omni/Linux infrastructure under the Omarchy project
    Engagement
    517 likes · 23 retweets · 22 replies
    Provenance
    Thread · Primary source
  3. 3

    François Chollet on benchmark claims and evaluation integrity

    X François Chollet — François Chollet created Keras and the ARC (Abstraction and Reasoning Corpus) benchmark at Google. He's one of the most persistent voices arguing for genuine reasoning evaluation over pattern-matching benchmarks.

    Chollet pushes back against teams claiming high benchmark scores without actual evaluation: 'you should not claim you have scored 100% on a benchmark if your model has never been evaluated on it.'

    x.com/fchollet/status/2094423666138427501 →
    Details
    Excerpt
    Chollet pushes back against teams claiming high benchmark scores without actual evaluation: 'you should not claim you have scored 100% on a benchmark if your model has never been evaluated on it.'
    Context
    Benchmark integrity is a growing problem. If teams can claim scores without being evaluated, the signal-to-noise ratio in model capability comparisons gets worse — and that affects procurement decisions, research direction, and public trust.
    Key points
    • Chollet objects to unverified benchmark claims, specifically calling out teams that haven't actually run their models on the benchmark
    • He notes there are plenty of private benchmarks available outside Kaggle for legitimate evaluation
    • The comment comes amid ongoing questions about whether some labs are claiming results from proxy evaluations rather than direct measurement
    Provenance
    Tweet · Primary source
  4. 4

    Apple's New Mac Line is Built Around Local AI. The Bet Is You'd Rather Own Than Rent.

    Video Nate B Jones, AI News & Strategy Daily — Nate B Jones has spent 20 years in tech; recently focuses on helping leaders use AI in their businesses. His channel provides structured analysis of hardware and strategy angles the main feed often skims over.

    Comprehensive breakdown of Apple's complete desktop refresh — Mac Mini with M6/M5 Pro, Mac Studio with M5 Max/Ultra up to 512 GB unified memory. The central question: owning local compute versus renting intelligence fro…

    www.youtube.com/watch?v=1lO8aNSLPJc →
    Details
    Excerpt
    Comprehensive breakdown of Apple's complete desktop refresh — Mac Mini with M6/M5 Pro, Mac Studio with M5 Max/Ultra up to 512 GB unified memory. The central question: owning local compute versus renting intelligence from frontier labs.
    Context
    Apple is making its largest-ever local AI bet at the exact moment the most capable agents are moving toward cloud-based persistent computers. It forces a choice about which future of personal computing you're building for, and it's not an either/or market — there's room for Apple, Nvidia, OpenAI, and others to coexist in different layers.
    Key points
    • Apple refreshed the entire desktop line around local AI: M6 Mac Mini (base), M5 Pro/Max/Ultra tiers up to 512 GB unified memory at $5,500+
    • Apple put the new M6 chip only at the bottom of the line while keeping Mini and Studio on M5 — signaling urgency about memory availability over chip-family neatness
    • The bet: people prefer paying once for hardware plus electricity rather than unknown monthly token costs; 512 GB won't hold every frontier model but local models can handle huge shares of real work
    • Apple explicitly selling the Mac as 'a computer where an AI agent can live' — the first time they've been this direct about it
    • Nvidia's DGX Spark competes as a dedicated AI appliance, but most people want the computer to be the AI machine
    Provenance
    Video · Supporting source
  5. 5

    Google WikiSkill paper on persistent agents and skills frameworks

    X Elvis (oskar.van.s) — Oskar van der Linde, Google researcher working on agent architecture and knowledge systems.

    Oskar highlights Google's WikiSkill paper as showing the effectiveness of persistent agents backed by knowledge bases and a skill wiki — extending Andrej Karpathy's LLM Wiki concept into a concrete operational framework.

    x.com/omarsar0/status/2094432587821482036 →
    Details
    Excerpt
    Oskar highlights Google's WikiSkill paper as showing the effectiveness of persistent agents backed by knowledge bases and a skill wiki — extending Andrej Karpathy's LLM Wiki concept into a concrete operational framework.
    Context
    If persistent agents can maintain and grow skill libraries over time, that changes the unit of value from model capability to agent persistence — a subtle but meaningful shift in how we think about building with AI.
    Key points
    • Google's WikiSkill demonstrates a framework for how agents can tap into a wiki of skills rather than relying solely on the model's training data
    • Persistent agents with knowledge bases outperform context-only approaches in the paper's benchmarks
    • Builds on Karpathy's earlier LLM Wiki idea but moves from concept to concrete architecture
    Provenance
    Tweet · Primary source
  6. 6

    Omacom Foundation raises to $12M

    Thread DHH

    Omacom Foundation has raised another $2 million from @xdanger and @brian_armstrong for now a combined TWELVE MILLION DOLLARS from our Founding Patrons.

    x.com/dhh/status/2094412164220031385 →
    Details
    Cited text
    Omacom Foundation has raised another $2 million from @xdanger and @brian_armstrong for now a combined TWELVE MILLION DOLLARS from our Founding Patrons.
    Context
    A Linux desktop foundation raising this much capital is unusual — it signals that founders see the agent-layer operating system as the next meaningful platform opportunity. The fact that both crypto and web-dev capitals are flowing into the same bet is worth noting.
    Key points
    • Omacom Foundation reached $12M in founding patron funding
    • Brian Armstrong (Coinbase CEO) and Yunjie Dai (@xdanger) each contributed to the latest round
    • 1Password and 37signals became Distinguished Corporate Patrons at $100K/year over three years
    • DHH frames it as a 'malleable OS' with perpetual fundraising model
    Engagement
    1271 likes · 61 retweets
    Provenance
    Thread · Primary source
  7. 7

    1Password and 37signals become Distinguished Corporate Patrons

    Article DHH

    The corporate patronage model is interesting because it avoids the traditional VC dynamic — no equity, no board seat. It's more like sponsorship with a mission statement.

    omarchy.org/news/2026/08/1password-and-37si… →
    Details
    Context
    The corporate patronage model is interesting because it avoids the traditional VC dynamic — no equity, no board seat. It's more like sponsorship with a mission statement.
    Key points
    • 1Password and 37signals commit $100K/year for 3 years each
    • Corporate money buys the same recognition as individual money per DHH's framing
    • 37signals already fully migrated their technical team to Omarchy
    Provenance
    Article · Supporting source
  8. 8

    Agent memory as a file format

    Article Cal Paterson

    Paterson's argument hits at a real tension in agent infrastructure: are we building tooling that scales with model capability, or locking people into complex extraction pipelines? The file-format approach means agents c…

    calpaterson.com/memoryfields.html →
    Details
    Context
    Paterson's argument hits at a real tension in agent infrastructure: are we building tooling that scales with model capability, or locking people into complex extraction pipelines? The file-format approach means agents can use any access pattern — bash, perl, sqlite queries.
    Key points
    • Proposes 'memoryfield' — a markdown + SQLite zip file format for agent memory
    • Argues against complex RAG pipelines, graph databases, and multi-stage extraction
    • Uses semantic search via vector embeddings as the retrieval mechanism
    • Design principle: fewer moving parts, let agents read raw prose
    Provenance
    Article · Supporting source
  9. 9

    Chollet on ARC 3 eval methodology

    Thread François Chollet

    The other eval process, the one used for frontier model APIs and that we run ourselves when one of our partners asks, is for frontier model APIs.

    x.com/fchollet/status/2094416334952255586 →
    Details
    Cited text
    The other eval process, the one used for frontier model APIs and that we run ourselves when one of our partners asks, is for frontier model APIs.
    Context
    This reveals a structural issue in benchmark governance — who controls the eval and how transparent is it? When Chollet says there's a separate eval for frontier model APIs, it suggests two different measurement systems coexist.
    Key points
    • Chollet distinguishes between Kaggle's private-set competition and a separate eval for frontier model APIs
    • JFPuget challenged whether the competition should allow semi-private data testing (possible in ARC-AGI 2, removed in ARC-AGI 3)
    • Chollet pushed back: don't claim scores on benchmarks you haven't evaluated against
    Engagement
    4 likes · 0 retweets
    Provenance
    Thread · Primary source
  10. 10

    JFPuget on ARC 3 benchmark limitations

    Thread JFPuget

    You refuse to do the same for arc agi3 this year. Fine, your call, your benchmark, your business.

    x.com/JFPuget/status/2094421787937304906 →
    Details
    Cited text
    You refuse to do the same for arc agi3 this year. Fine, your call, your benchmark, your business.
    Context
    The benchmark rules changed between years in ways that affect competitive strategies, and the foundation doesn't seem willing to accommodate previous testing approaches. This is the kind of invisible infrastructure decision that matters more than most public announcements.
    Key points
    • ARC-AGI 2 allowed harness evaluation with frontier models (Yohan Land scored 72.9%)
    • ARC-AGI 3 removed the ability to test on semi-private data
    • JFPuget is frustrated that prior methodology changes can't be challenged retroactively
    Provenance
    Thread · Primary source
  11. 11

    A CVE dispute

    Article Daniel Stenberg

    Daniel makes a practical point: CVEs are not just technical designations, they're distributed-work triggers. When a CVE gets assigned, it automatically fires off patches and updates across the entire ecosystem. Not all…

    daniel.haxx.se/blog/2026/06/24/a-cve-dispute →
    Details
    Context
    Daniel makes a practical point: CVEs are not just technical designations, they're distributed-work triggers. When a CVE gets assigned, it automatically fires off patches and updates across the entire ecosystem. Not all bugs deserve that signal.
    Key points
    • curl became a CNA and had its first-ever CVE dispute with MITRE
    • The disputed issue: a wildcard hostname-checking bug that required highly unlikely conditions to exploit
    • MITRE TL-Root ruled against assigning a CVE, agreeing with curl's assessment
    • Every CVE has ecosystem cost — patch deployment across ~30 billion libcurl instances
    Provenance
    Article · Supporting source
  12. 12

    Codex Tool Reference

    Article Simon Willison

    Simon's reference captures what 'full-featured' agent tooling looks like today: a massive surface area of capabilities that any model session could potentially access. The browser control skill in particular is where th…

    codex-tool-reference.simonw.chatgpt.site →
    Details
    Context
    Simon's reference captures what 'full-featured' agent tooling looks like today: a massive surface area of capabilities that any model session could potentially access. The browser control skill in particular is where the real agentic work happens.
    Key points
    • 232 tool interfaces and 44 complete skills in a single Codex Work session
    • Skills cover browser control, document generation, data analytics, presentations, and more
    • Tool declarations include TypeScript-style type signatures with sandbox_permissions, max_output_tokens, etc.
    Provenance
    Article · Supporting source