SemiAnalysis reports that OpenAI’s Broadcom-built ASIC delivered 1.5–1.9 times the performance per watt over Nvidia, AMD, and Google hardware on DeepSeek R1 and Kimi. Its measured latency advantage ranged from 1.7 to 3.6 times. The third-party benchmarks give buyers the first public numbers for a chip that OpenAI announced in June.
Read source◆ Braid Daily · 2026-08-26
OpenAI’s Jalapeño challenges the GPU rack
Third-party benchmarks put OpenAI’s custom ASIC ahead on performance per watt and latency across two open models.
The lead
1Chinese models become products
4Z.ai identifies Ox Alpha and schedules the weights
Bloomberg via Techmeme
Z.ai confirmed that the model leading OpenRouter’s usage chart is a new GLM iteration and said it would release the weights tonight. The attribution ends the model’s anonymous run and gives local operators a concrete download window.
Read sourceAlibaba releases Qwen3.8-Flash
Bloomberg via Techmeme
Alibaba’s new open-weight model has 125 billion parameters and uses its next-generation Qwen 4 architecture. The available item doesn’t include an independent evaluation of Alibaba’s claim that it rivals Opus 4.6 and V4-Flash.
Read sourceMoonshot seeks a cut of hyperscaler Kimi K3 revenue
Reuters via Techmeme
Reuters sources say Moonshot is in early talks with Microsoft, Amazon, and Google about hosting Kimi K3, with a requested revenue share of as much as 30%. The talks would turn distribution of an open-weights model into recurring platform revenue for its lab.
Read sourceDeepSeek’s revenue rises alongside a larger loss
The Information via Techmeme
The Information’s sources put DeepSeek’s revenue at $70.7 million for the first seven months of 2026, about ten times its full-year 2025 revenue. The company posted a $106 million net loss over the same seven-month period, compared with a $139 million loss in 2025.
Read sourceMeasuring agents at work
4Meta tested deep cuts, then rejected the agent math
Reuters via Techmeme
Internal documents reviewed by Reuters show that Meta explored reducing many teams by about 60% as part of an AI-native plan. The company pulled back after employee resistance and internal data showed that the agents weren’t effective enough.
Read sourceloveholidays ties Codex adoption to operating results
OpenAI
loveholidays says AI-assisted coding grew from nearly zero to about 80% of code production within a year. Over that period, deployment frequency rose 73% at stable headcount, data-platform modifications doubled, and support tickets fell by half.
Read sourcePerplexity moves its agent loop onto local Nvidia machines
VentureBeat via Techmeme
Portable Computer runs Perplexity’s agent platform on-device, starting with Nvidia DGX Spark and RTX Linux PCs. Perplexity describes the local execution path as carrying no token charges, which changes both the privacy boundary and the recurring inference bill.
Read sourceParitok compresses coding context to about one quarter
arXiv
Paritok-4B compresses SWE-bench Lite agent context to 25.7% of its original size while retaining 86.5% of uncompressed single-shot solve quality, according to the authors. Its 264 MB Apache 2.0 adapter, training data, and evaluation scripts are open for testing in other harnesses.
Read sourceWho controls and funds the compute
4Two labs could consume most marginal compute by 2028
Dwarkesh Patel
Dylan Patel projects that OpenAI and Anthropic’s share of marginal compute capacity will rise from about 30% now to 40–50% next year, then reach 70–80% by 2028. He estimates annual infrastructure costs at $10–15 million per megawatt. Frontier inference can generate as much as $50 million per megawatt, according to his estimate.
Read sourceLenders try to contain data-center exposure
Financial Times via Techmeme
Major lenders are financing, insuring, and underwriting AI data centers as a new asset class while trying to limit their own exposure. The financing constraints sit directly beneath the trillion-dollar infrastructure forecasts.
Read sourceThe EPA proposal removes public notice from some air permits
Tom’s Hardware
A proposed EPA change would allow some data-center air-pollution permits to proceed without public notice or input. The procedural change addresses local resistance by narrowing where communities can object during permitting.
Read sourceAustralia adds state carveouts to renewable-power rules
The Guardian
Australia’s national cabinet opened a path for Queensland, the Northern Territory, and other jurisdictions with state-owned power systems to receive flexibility under national data-center standards. The change retreats from a requirement that new AI data centers run entirely on renewable energy.
Read sourceA separate channel for untrusted text
1Semantic Overlays mark spans outside the token stream
arXiv
Semantic Overlays apply learned adapters to selected prefill positions, giving a frozen model a non-text channel for identifying untrusted spans. The author reports that TensorTrust attack success fell from 34.8% to 6.6%. All four PIArena attack families reached zero compliance, while the system exactly copied 92.5% of marked text.
Read sourceCompanion episode