Archive BRAID DAILY
Grok 4.5 puts coding agents back under test
Subscribe

Braid Daily · 2026-07-09

Grok 4.5 puts coding agents back under test

Today’s issue treats Grok 4.5 as a coding-and-agents launch, then checks the claims against practical tests and benchmark audits.

Dark editorial cover image showing coding-agent systems, model evaluation traces, and compute hardware as one restrained technical scene.
Braid Daily cover art for the July 9, 2026 issue.

The lead

1

SpaceXAI announced Grok 4.5 as a frontier model for coding and agents. That puts today’s launch in direct conversation with software teams’ toolchains, not just the general chatbot cycle.

Read source

Models In The Toolchain

4

OpenAI demos full-duplex ChatGPT Voice

OpenAI on YouTube

OpenAI’s GPT Live 1 demo broadens the day beyond coding. The candidate set points to full-duplex audio, an intelligence picker, and delegated search or reasoning as the product details to inspect.

Read source

Reading Coding-Agent Scores

3

AgentLens scores the whole agent trajectory

arXiv

AgentLens evaluates code agents across the execution trajectory, not only the final pass or fail. That distinction matters when an agent reaches the right answer through brittle steps a production workflow couldn’t tolerate.

Read source

Compute, Control Planes, And Policy Pressure

5

Positron raises fresh AI-chip capital

Techmeme

Positron’s funding news belongs in the hardware column because inference cost and power use remain open constraints. The candidate set treats it as one funding marker among several, not proof of a changed chip market.

Read source

AWS introduces Claude Apps Gateway

AWS Machine Learning Blog

AWS is positioning Claude Apps Gateway as a self-hosted control plane for enterprise access and cost management. Operators get another place to enforce usage policy before requests reach the model provider.

Read source

AI companies move election spending into view

CNBC

CNBC’s reporting moves the policy fight into campaign spending and AI legislation. It sits next to the compute story because release authority, procurement, and regulation now move through political channels as well as product channels.

Read source

Companion episode

Coding Models Meet Their Test Bench

· 00:26:22

Comparative evidence beats launch energy here: which coding agents survive large repositories, audited benchmarks, and enterprise control planes together. Today’s items give buyers more places to test that claim.