Archive BRAIXD
Astra, agents, and the PSP model / DISPATCH 118
PDF RSS

Dispatch 118 · 2026-09-05 Braixd

Astra, agents, and the PSP model

/ 00:17:08 / 14 sources

“The launch was so chaotic that ChatGPT, Claude, Grok, and Cursor all went down at the same time. The Occam's razor explanation: Azure had an outage. The funnier theory: Astra killed its competition on day one.”

— Seln Oriax, today's narration

OpenAI launched GPT-6 Astra on Friday with a launch so messy it took down its own competitors. Meanwhile, DHH sees local agents converging into something "out of the box," someone is running a 90M parameter LLM on a 22-year-old Sony PSP at half a token per second, and prediction market theorist Robin Hanson asks what fraction of 2028 voters will consult an AI — and actually follow its advice.

Chapters

  1. 00:00:04 The Astra launch that took everyone down
  2. 00:05:46 Local agents and the PSP model — the edge is alive
  3. 00:10:17 How many voters will ask an LLM — and follow it?
  4. 00:14:22 Weekend wrap — what the local pass reveals

Sources

14 cited
  1. 1

    You can now run a 90M conversational LLM on the Sony PSP (hardware from 2004)

    Article liright / LocalLLaMA

    I wanted to see what the PSP can theoretically handle and I got my answer - a 90M model is about the max it can do without atrocious inference speeds. It's running around 0.5 - 0.6 tokens per second, which is very slow,…

    www.reddit.com/r/LocalLLaMA/comments/1w78zt… →
    Details
    Cited text
    I wanted to see what the PSP can theoretically handle and I got my answer - a 90M model is about the max it can do without atrocious inference speeds. It's running around 0.5 - 0.6 tokens per second, which is very slow, but it's useable. Maybe 1-3 minutes for a reply.
    Key points
    • 90M parameter model running on Sony PSP (2004 hardware)
    • Inference speed: ~0.5-0.6 tokens/second, 1-3 minute replies
    • Model can generate 'crappy poems, short stories, write non-functional code'
    • GitHub: thatblend/LLMPSP
    Provenance
    Article · Supporting source
  2. 2

    Did OpenAI actually build AGI? GPT-6 Astra first look — Fireship Code Report

    Source Fireship / Jeff Delaney

    Fireship covers the week's model cascade: Anthropic Fable/Mythos 5.1 Tuesday, Meta MuSpark 1.3 Wednesday, OpenAI GPT-6 Astra Friday. Covers the concurrent outage of ChatGPT/Claude/Grok/Cursor, messy rollout drama, bench…

    www.youtube.com/watch?v=FluKUJyeYD8 →
    Details
    Excerpt
    Fireship covers the week's model cascade: Anthropic Fable/Mythos 5.1 Tuesday, Meta MuSpark 1.3 Wednesday, OpenAI GPT-6 Astra Friday. Covers the concurrent outage of ChatGPT/Claude/Grok/Cursor, messy rollout drama, benchmark numbers (OSWorld 73%, Arc AGI 3 99%), and the AI Intelligence Index score of 61 matching GPT-5.6 Soul but trailing Fable 5.1 by 5 points.
    Key points
    • ChatGPT, Claude, Grok, Cursor all went down simultaneously right before Astra launch
    • Astra scored 73% on OSWorld (desktop task benchmark) in ~40 min vs Soul's 65% at 75 min
    • First model to hit OpenAI's 'critical cyber threshold' — can find and exploit zero days autonomously
    • AI Intelligence Index: Astra 61, GPT-5.6 Soul 61, Fable 5.1 66
    Provenance
    Source · Background source
  3. 3

    Introducing GPT-6 Astra for developers — OpenAI official demo

    Source OpenAI / Charlie Guo, Developer Experience Engineer

    Charlie Guo walks through GPT-6 Astra's new capabilities: improved computer use via screenshots, better creative/knowledge work outputs, 3D model building, and the new async tool calling + steering features in the Respo…

    www.youtube.com/watch?v=bOC3DisEOfg →
    Details
    Excerpt
    Charlie Guo walks through GPT-6 Astra's new capabilities: improved computer use via screenshots, better creative/knowledge work outputs, 3D model building, and the new async tool calling + steering features in the Responses API. Available in ChatGPT, Codex, and the API.
    Key points
    • Computer use uses screenshots to track app state while keeping the app in background
    • Async tool calling lets model work on other task parts while a tool call runs
    • Steering lets you change direction mid-response without canceling the running tool
    • Pricing: $10M input / $50M output tokens, same as Fable 5.1
    Provenance
    Source · Background source
  4. 4

    .gitignore everything by default — Alex Pliutau

    Article Alex Pliutau

    A proposal to flip the conventional .gitignore approach: instead of allowing everything and selectively ignoring, ignore everything and only allow specific files. Uses the * followed by ! negation syntax to create a whi…

    packagemain.tech/p/gitignore-everything-by-… →
    Details
    Excerpt
    A proposal to flip the conventional .gitignore approach: instead of allowing everything and selectively ignoring, ignore everything and only allow specific files. Uses the * followed by ! negation syntax to create a whitelist-based git tracking strategy. The author notes modern projects have so much local junk (agentic docs, subfolders) that starting from 'ignore everything' may feel easier.
    Key points
    • Flips .gitignore philosophy: ignore * by default, ! allow what you need
    • Uses git's built-in negation syntax (* then ! for exceptions)
    • Author notes modern projects generate so much local junk that whitelist approach may be easier
    • References typescript-go's 207-line .gitignore as evidence of current repo bloat
    Provenance
    Article · Supporting source
  5. 5

    DHH on local agents, preconfigured

    X DHH / David Heinemeier Hansson

    This is so cool. Local agents, preconfigured. I know @0xsero is also cooking here. We're going to end up with something amazing out of the box very soon!

    x.com/dhh/status/2096233810224386241 →
    Details
    Cited text
    This is so cool. Local agents, preconfigured. I know @0xsero is also cooking here. We're going to end up with something amazing out of the box very soon!
    Key points
    • DHH sees a convergence moment around preconfigured local agents
    • Mentions @0xSero as someone building in this space
    • Expects 'something amazing out of the box' — suggesting consumer-ready tooling is close
    Engagement
    79 likes · 2 retweets · 10 replies
    Provenance
    Tweet · Primary source
  6. 6

    Robin Hanson on LLM voting advice — two-part question

    X Robin Hanson / George Mason University economics professor, prediction market theorist

    What % of those who vote in the 2028 US presidential election will ask a LLM for a recommendation? What % of those who ask an LLM how to vote in the 2028 US presidential election will NOT do what the LLM advises?

    x.com/robinhanson/status/2096249543230627991 →
    Details
    Cited text
    What % of those who vote in the 2028 US presidential election will ask a LLM for a recommendation? What % of those who ask an LLM how to vote in the 2028 US presidential election will NOT do what the LLM advises?
    Key points
    • Hanson frames both adoption and compliance as open questions
    • The two-part structure reveals skepticism about whether people will actually follow AI advice
    • Published Saturday Sept 5, 2026 — four years before the election
    Provenance
    Tweet · Primary source
  7. 7

    Anatoly Karlin's reply to Hanson on AI voting advice

    X Anatoly Karlin / economist and political commentator

    I think the AI will be smarter than me by 2028. I will not only ask but implement as it instructs.

    x.com/akarlin/status/2096254134630662578 →
    Details
    Cited text
    I think the AI will be smarter than me by 2028. I will not only ask but implement as it instructs.
    Key points
    • Karlin's answer is strikingly different from Hanson's framing — he expects compliance, not skepticism
    • The word 'implement' suggests acting on the advice, not just considering it
    • Published Saturday Sept 5, 2026
    Provenance
    Tweet · Primary source
  8. 8

    GPT-6 Astra with Ben Davis (OpenAI)

    Video OpenAI — Ben Davis is a puzzle designer and competitive hacker who organized DEF CON puzzles for years.

    Ben Davis tests GPT-6 Astra on DEF CON puzzle challenges, describing parallel research branches and sub-agent swarm workflows that let the model keep itself on track through complex multi-step reasoning.

    www.youtube.com/watch?v=B-jjnydci50 →
    Details
    Excerpt
    Ben Davis tests GPT-6 Astra on DEF CON puzzle challenges, describing parallel research branches and sub-agent swarm workflows that let the model keep itself on track through complex multi-step reasoning.
    Context
    The DEF CON puzzle results are one of the few publicly observable benchmarks that aren't just another math leaderboard. The multi-agent architecture is OpenAI's explicit answer to long-horizon reasoning.
    Key points
    • Solved three puzzles that DEF CON organizers hadn't solved, plus one unsolved by anyone in the world
    • Model uses ~10 parallel agent slots to test its own theories
    • Key differentiator: better at keeping itself on track compared to prior models
    • Sub-agent/swarm workflows described as a new capability worth trying
    Provenance
    Video · Supporting source
  9. 9

    Solaris: an interface world model

    Source Cristóbal Valenzuela

    Solaris is described as an interactive, real-time video model that creates and renders interfaces, with Nandan Priyadarshi calling it the 'missing middle' between text agents and click-the-UI workflows.

    x.com/c_valenzuelab/status/2096208484714823… →
    Details
    Excerpt
    Solaris is described as an interactive, real-time video model that creates and renders interfaces, with Nandan Priyadarshi calling it the 'missing middle' between text agents and click-the-UI workflows.
    Context
    If interface generation moves from static screenshots to live rendered response, the whole text-to-action pipeline changes. The hallucinated button problem gets solved differently.
    Key points
    • Interface world model: generates UI in real-time video
    • Responds to user interaction as it renders
    • Solves the gap between agent output and actual tool use
    Provenance
    Source · Background source
  10. 10

    Robin Hanson on LLM jagged abilities

    Source Robin Hanson — Robin Hanson is an economics professor at George Mason University known for his work on signaling, consensus, and forecasting.

    'HLLMs abilities are jagged, and seem to be missing some big chunks of what really matters in human intelligence.' Quoting his own earlier observation about models going through the motions without true understanding.

    x.com/robinhanson/status/2096250777928859978 →
    Details
    Excerpt
    'HLLMs abilities are jagged, and seem to be missing some big chunks of what really matters in human intelligence.' Quoting his own earlier observation about models going through the motions without true understanding.
    Context
    Hanson has been tracking this since early. His 'jagged' framing is useful because it predicts where models will fail even as they improve on the benchmarks that got them there.
    Key points
    • LLM capabilities are uneven — strong in some domains, missing fundamentals in others
    • The gap isn't just performance; it's structural: lack of step-back reasoning
    • This is a repeat observation, not a new finding, but persists even with frontier models
    Provenance
    Source · Background source
  11. 11

    Claude's new system prompt really doesn't want to reproduce song lyrics

    Article Simon Willison — Simon Willison has tracked Anthropic's system prompts via a git repo he maintains. He also built an automation pipeline using GPT-5.6 Luna to summarize prompt diffs.

    Anthropic updated Claude Fable 5.1's system prompt with heavy new copyright sections: no reproducing song lyrics, poems, or book passages; no generating images of copyrighted characters; drug guidance reframed with harm…

    simonwillison.net/2026/Sep/2/claudes-new-sy… →
    Details
    Excerpt
    Anthropic updated Claude Fable 5.1's system prompt with heavy new copyright sections: no reproducing song lyrics, poems, or book passages; no generating images of copyrighted characters; drug guidance reframed with harm-reduction URLs; and a new stance against being submissive to abusive users.
    Context
    The timing around the Sony/Warner Chappell lawsuit is notable. But the more interesting shift is behavioral: Claude is now instructed not to be submissive when users are abusive, replacing the old warning-and-end protocol with 'accountability without self-abasement.'
    Key points
    • New section explicitly banning song lyric reproduction (any amount) with persistent refusal
    • Image copyright now extends to code-generated art (SVG, canvas, CSS mockups)
    • Claude drops the anti-dependency rules — no longer told not to thank users or invite continued conversation
    • Drug guidance now includes URLs to dancesafe.org, tripsit.me, and psychonautwiki.org — first time Claude's system prompt included non-Anthropic URLs
    Provenance
    Article · Supporting source
  12. 12

    There's No Limit to How Bad Code Can Get

    Article Zach Kehs

    A technical debt essay arguing that metaphors like 'sinking ship' are misleading because software has no natural collapse threshold — a business dies long before the code hits any hypothetical floor. Technical debt has…

    zachkehs.com/blog/theres_no_limit_to_how_ba… →
    Details
    Excerpt
    A technical debt essay arguing that metaphors like 'sinking ship' are misleading because software has no natural collapse threshold — a business dies long before the code hits any hypothetical floor. Technical debt has no bankruptcy, no clean reset.
    Context
    The piece has quiet relevance to the AI tooling debate: as more code gets generated, the question isn't whether tools can write good code, but what happens when the generated code is already buried under layers of previous generations with no reset option.
    Key points
    • Code can always get worse: new layers of indirection, performance degradation — there's no physical constraint like a collapsing building
    • 'Sinking ship' metaphor implies an end that doesn't exist in software
    • Mega-corporations handle it via side-channels (separate teams building disconnected systems) rather than resets
    Provenance
    Article · Supporting source
  13. 13

    Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries

    Article Amine Dirhoussi et al. (HuggingFace TRL team) — Amine Dirhoussi, Quentin Gallouédec, Kashif Rasul, Lewis Tunstall, Edward Beeching, Albert Villanova del Moral, Nouamane Tazi, Leandro von Werra, and Sergio Paniego from the HuggingFace TRL team.

    A detailed survey of 16 open-source async RL libraries showing convergence on disaggregated inference/training architectures. On a single H100, a 32B model generating 32K-token rollouts takes ~3.7 hours — making synchro…

    huggingface.co/blog/async-rl-training-lands… →
    Details
    Excerpt
    A detailed survey of 16 open-source async RL libraries showing convergence on disaggregated inference/training architectures. On a single H100, a 32B model generating 32K-token rollouts takes ~3.7 hours — making synchronous training impossible at scale.
    Context
    This is what happens when you take the Astra-level claims seriously at production scale: the infrastructure layer becomes the actual constraint. The article's number — 3.7 hours of generation per training step — is the kind of detail that rarely makes it into launch coverage.
    Key points
    • 16 libraries surveyed all converged on the same architecture: separate GPU pools for inference and training
    • Generation dominates wall-clock time: 512 rollouts of 2K tokens = 14 minutes on one H100; 32K tokens = ~3.7 hours
    • Ray dominates orchestration (8/16), NCCL broadcast is default weight sync, LoRA support is sparse, distributed MoE is the emerging differentiator
    Provenance
    Article · Supporting source
  14. 14

    Paul Graham on lab motivation

    Source Paul Graham — Paul Graham is a co-founder of Y Combinator, essayist, and long-time observer of the technology industry.

    'He might be right, but I think the two cases are different, if only because individual employees at the labs are genuinely curious about open problems.'

    x.com/paulg/status/2096230913688371332 →
    Details
    Excerpt
    'He might be right, but I think the two cases are different, if only because individual employees at the labs are genuinely curious about open problems.'
    Context
    Small tweet, but it lands between the Astra launch claims and Hanson's persistent skepticism about what models actually understand. Graham is pointing to something about why labs keep pushing forward even when the gains seem incremental.
    Key points
    • Graham challenges a prediction about lab motivation
    • Argues that individual researchers' genuine curiosity is a key differentiator from purely commercial incentives
    Provenance
    Source · Background source