Archive BRAID DAILY
OpenAI’s Jalapeño challenges the GPU rack
Subscribe

Braid Daily · 2026-08-26

OpenAI’s Jalapeño challenges the GPU rack

Third-party benchmarks put OpenAI’s custom ASIC ahead on performance per watt and latency across two open models.

A signal-yellow AI accelerator glows among dark data-center racks.

The lead

1

SemiAnalysis reports that OpenAI’s Broadcom-built ASIC delivered 1.5–1.9 times the performance per watt over Nvidia, AMD, and Google hardware on DeepSeek R1 and Kimi. Its measured latency advantage ranged from 1.7 to 3.6 times. The third-party benchmarks give buyers the first public numbers for a chip that OpenAI announced in June.

Read source

Chinese models become products

4

Z.ai identifies Ox Alpha and schedules the weights

Bloomberg via Techmeme

Z.ai confirmed that the model leading OpenRouter’s usage chart is a new GLM iteration and said it would release the weights tonight. The attribution ends the model’s anonymous run and gives local operators a concrete download window.

Read source

Alibaba releases Qwen3.8-Flash

Bloomberg via Techmeme

Alibaba’s new open-weight model has 125 billion parameters and uses its next-generation Qwen 4 architecture. The available item doesn’t include an independent evaluation of Alibaba’s claim that it rivals Opus 4.6 and V4-Flash.

Read source

Moonshot seeks a cut of hyperscaler Kimi K3 revenue

Reuters via Techmeme

Reuters sources say Moonshot is in early talks with Microsoft, Amazon, and Google about hosting Kimi K3, with a requested revenue share of as much as 30%. The talks would turn distribution of an open-weights model into recurring platform revenue for its lab.

Read source

DeepSeek’s revenue rises alongside a larger loss

The Information via Techmeme

The Information’s sources put DeepSeek’s revenue at $70.7 million for the first seven months of 2026, about ten times its full-year 2025 revenue. The company posted a $106 million net loss over the same seven-month period, compared with a $139 million loss in 2025.

Read source

Measuring agents at work

4

Meta tested deep cuts, then rejected the agent math

Reuters via Techmeme

Internal documents reviewed by Reuters show that Meta explored reducing many teams by about 60% as part of an AI-native plan. The company pulled back after employee resistance and internal data showed that the agents weren’t effective enough.

Read source

loveholidays ties Codex adoption to operating results

OpenAI

loveholidays says AI-assisted coding grew from nearly zero to about 80% of code production within a year. Over that period, deployment frequency rose 73% at stable headcount, data-platform modifications doubled, and support tickets fell by half.

Read source

Perplexity moves its agent loop onto local Nvidia machines

VentureBeat via Techmeme

Portable Computer runs Perplexity’s agent platform on-device, starting with Nvidia DGX Spark and RTX Linux PCs. Perplexity describes the local execution path as carrying no token charges, which changes both the privacy boundary and the recurring inference bill.

Read source

Paritok compresses coding context to about one quarter

arXiv

Paritok-4B compresses SWE-bench Lite agent context to 25.7% of its original size while retaining 86.5% of uncompressed single-shot solve quality, according to the authors. Its 264 MB Apache 2.0 adapter, training data, and evaluation scripts are open for testing in other harnesses.

Read source

Who controls and funds the compute

4

Two labs could consume most marginal compute by 2028

Dwarkesh Patel

Dylan Patel projects that OpenAI and Anthropic’s share of marginal compute capacity will rise from about 30% now to 40–50% next year, then reach 70–80% by 2028. He estimates annual infrastructure costs at $10–15 million per megawatt. Frontier inference can generate as much as $50 million per megawatt, according to his estimate.

Read source

Lenders try to contain data-center exposure

Financial Times via Techmeme

Major lenders are financing, insuring, and underwriting AI data centers as a new asset class while trying to limit their own exposure. The financing constraints sit directly beneath the trillion-dollar infrastructure forecasts.

Read source

Australia adds state carveouts to renewable-power rules

The Guardian

Australia’s national cabinet opened a path for Queensland, the Northern Territory, and other jurisdictions with state-owned power systems to receive flexibility under national data-center standards. The change retreats from a requirement that new AI data centers run entirely on renewable energy.

Read source

A separate channel for untrusted text

1

Semantic Overlays mark spans outside the token stream

arXiv

Semantic Overlays apply learned adapters to selected prefill positions, giving a frozen model a non-text channel for identifying untrusted spans. The author reports that TensorTrust attack success fell from 34.8% to 6.6%. All four PIArena attack families reached zero compliance, while the system exactly copied 92.5% of marked text.

Read source

Companion episode

Whose numbers are these

· 00:26:43