Archive BRAID DAILY
Astra's benchmark weekend has a moving baseline
Subscribe

Braid Daily · 2026-09-06

Astra's benchmark weekend has a moving baseline

External tests look strong, but OpenAI's published metrics changed after launch. This issue separates the evidence by source.

A dark editorial illustration of shifting calibration plates observed by external probes.
Astra reached its first weekend with strong outside results and vendor metrics that were still changing.

The lead

1

GPT-6 Astra's first weekend produced strong external results in security, code, and robot control. Fortune reports that OpenAI also changed several evaluation metrics after launch, so the origin and timing of each number matter.

Read source
A diagram separating OpenAI metrics, external Astra results, and one unverified jailbreak report.
The source of each claim matters while OpenAI is revising its published metrics.

Astra's first external tests

4

Astra scores 95% on a robot-control task

r/singularity

The published comparison puts Astra at 95% on a robot-control task. Fable 5.1 scored 40%. This is an application-specific result, but it tests physical control rather than another language benchmark.

Read source

Incident disclosure meets multi-agent coordination

2

Choosing the harness and its supporting systems

3

NInfer, llama.cpp, and vLLM on one RTX 5090

r/LocalLLaMA

A community comparison runs Qwen3.8-27B NVFP4 across three inference engines on an RTX 5090. The figures come from one person's rig and workload, but the setup is concrete enough to inform a local deployment shortlist.

Read source

Costs measured outside the model

3

Companion episode

Someone Else's Stopwatch

· 00:22:56