Fireship surveys several recent AI-assisted math results, including counterexamples, new sphere-packing bounds, and work checked by mathematicians. Formal verification in Lean is becoming part of the result rather than an optional appendix.
Read source◆ Braid Daily · 2026-08-20
Machine-checked proof becomes the acceptance bar
AI-assisted mathematics is producing counterexamples and new bounds; formal verification is becoming part of the result.
The lead
1Shipping health AI without patient experiments
5Ufonia puts simulation before deployment
AI Engineer
Ufonia says its clinical voice agent has completed about 200,000 calls across 20 UK hospitals. Its release process simulates patients, scores hazardous scenarios with an automated judge validated against ten clinicians, and moves forward only after the simulation threshold is met.
Read sourceHinge Health moves critical decisions into code
AI Engineer
Hinge Health puts deterministic code ahead of the model for emergency escalation, intent routing, and identity verification. Its release process judges safety bugs by the worst plausible outcome and checks evaluator calibration before changing prompts.
Read sourceAnterior makes the audit trail the system of record
AI Engineer
Anterior describes an event ledger that records new actions without rewriting history, protected data stored as immutable blobs, and one action chain for clinicians and models. The ledger can replay the exact state used for an earlier decision without exposing raw patient data.
Read sourceHippocratic runs 31 clinical specialists in parallel
AI Engineer
Hippocratic says 31 specialist models support its clinical voice reasoning while four-bit quantization, speculative decoding, and key-value cache compression keep latency down. The company reports a 96 percent cache hit rate; the figures are vendor claims rather than independent benchmark results.
Read sourceAnterior generates patient records backward from labels
AI Engineer
Anterior samples a target label, derives a policy-driven reasoning path, and then generates the unstructured medical record. The company says about 90 percent of its datasets are synthetic and that clinicians identified synthetic records with 60 percent accuracy in blind evaluations.
Read sourceModels, agents, and inference
5xAI ships Grok Build
Elon Musk on X
xAI announced Grok Build alongside the first user reports of repository-level workflows. Its benchmark claims come from sources close to the release; the concrete evidence so far is that users are applying the system to existing repositories.
Read sourceClaude Managed Agents add domain allowlists
Claude Developers on X
Claude Managed Agents can now constrain web search and fetching to approved domains. This narrows the retrieval path used by poisoned skills and delayed payloads, the agent-security problem covered on Tuesday.
Read sourceOrnith-1.5 spans three model sizes
Ornith
Ornith’s release covers models with 9, 35, and 397 billion parameters. They focus on models that assemble and improve their own agent workflows, while the coding and agent results remain release claims pending independent reproduction.
Read sourceDeepSeek releases V4 Pro weights under MIT
Two Minute Papers
DeepSeek released V4 Pro weights under the MIT license while raising official hosted prices by 2.5 to 5 times. The reported training recipe distills more than ten specialist teachers into one model, with speculative decoding used to accelerate inference.
Read sourceDFlash2 moves from write-up to local reproduction
Inco
DFlash2 proposes parallel drafting for faster inference. Within roughly eight hours, a vLLM pull request and a consumer RTX 3090 reproduction followed, shortening the path from method to runnable implementation.
Read sourceControl and infrastructure
3OpenRouter confirms it is joining Stripe
OpenRouter
Following the reported deal covered on August 17, OpenRouter has confirmed in its own announcement that it is joining Stripe. The primary source turns a reported acquisition into an acknowledged platform combination.
Read sourceA bipartisan bill proposes an AI shutdown authority
Ted Lieu on X
Ted Lieu and Nathaniel Moran announced a bipartisan bill aimed at establishing shutdown authority for dangerous AI systems. Their same-day posts set out the proposal’s political and safety case.
Read sourceOpenAI introduces Private Safety Processing
OpenAI on X
OpenAI presents Private Safety Processing as a way to reconcile frontier safety detection with Zero Data Retention customers. Enterprise teams now have a concrete privacy guarantee and a set of limits to inspect.
Read sourceCompanion episode