Archive BRAIXD
Search is broken, models keep adapting, and agents learn from their own mistakes / DISPATCH 070
PDF RSS

Dispatch 070 · 2026-07-05

Search is broken, models keep adapting, and agents learn from their own mistakes

/ 00:05:56 / 6 sources

“AI could be wrong about half the time. That's not a benchmark result — it's a Monday morning search query.”

— Seln Oriax, today's narration

Two separate threads on the same problem today: a Wired fact-checker found AI-powered search engines inaccurate over 60% of the time, while on the other side, people are already building systems that adapt after deployment. The gap between what we can do and what we can trust is widening.

We also look at liquid models emerging from LiquidAI's LFM2.5 collection, Soheil Feizi's framework for agent continual learning, and a surprisingly practical use case for Claude Code with Slack MCP integration.

Chapters

  1. 00:00:04 The accuracy floor
  2. 00:02:03 Liquid models as a practical stack
  3. 00:03:33 Agents learning from their own mistakes
  4. 00:04:59 Where the quality sits

Sources

6 cited
  1. 1

    I'm a Professional Fact-Checker. AI Is Wrong More Often Than You Think

    Article Meghan Herbst, WIRED

    A WIRED fact-checker found AI search results wrong about a third of the time for basic lookup tasks, with studies putting inaccuracy between 45% and over 60%. The Tow Center study from March 2025 found more than 60 perc…

    www.wired.com/story/fact-checking-ai →
    Details
    Excerpt
    A WIRED fact-checker found AI search results wrong about a third of the time for basic lookup tasks, with studies putting inaccuracy between 45% and over 60%. The Tow Center study from March 2025 found more than 60 percent of AI-powered search responses were inaccurate.
    Key points
    • WIRED's fact-checker finds AI search results wrong ~1/3 of the time on basic lookup tasks
    • Tow Center March 2025 study: over 60% of AI search engine responses inaccurate
    • BBC study puts chatbot wrongness closer to 45%
    • SimpleQA benchmark: no model exceeded 50% accuracy on OpenAI's test set
    • Claude led RealFactBench at 73% accuracy; Grok wasn't assessed
    Provenance
    Article · Supporting source
  2. 2

    Robin Hanson on AI search inaccuracy

    X Robin Hanson

    x.com/robinhanson/status/2073732144623452578 →
    Details
    Engagement
    21 likes · 1 retweets · 7 replies
    Provenance
    Tweet · Primary source
  3. 3

    Zach Mueller on liquid models use case

    X Zach Mueller

    x.com/TheZachMueller/status/207373140814558… →
    Details
    Engagement
    4 likes · 2 retweets · 4 replies
    Provenance
    Tweet · Primary source
  4. 4

    LFM2.5 — a LiquidAI Collection

    Article LiquidAI

    A collection of post-trained and base LFM2.5 models ranging from 230M to 8B parameters, with multiple variants including thinking, instruct, Japanese, and GGUF checkpoints.

    huggingface.co/collections/LiquidAI/lfm25 →
    Details
    Excerpt
    A collection of post-trained and base LFM2.5 models ranging from 230M to 8B parameters, with multiple variants including thinking, instruct, Japanese, and GGUF checkpoints.
    Key points
    • LiquidAI released LFM2.5 model family: 230M, 350M, 1.2B, and 8B-A1B variants
    • Models span text generation with thinking and instruct modes
    • GGUF-compatible checkpoints available alongside standard formats
    Provenance
    Article · Supporting source
  5. 5

    Continual Learning for AI Agents: From Failures to Durable Improvements

    Video Soheil Feizi, RELAI / UMD

    Feizi presented a framework for converting production failures into testable, regression-aware improvements. Key insight: raw session logs plus feedback aren't enough—you need replayable learning environments to make fa…

    www.youtube.com/watch?v=2IxD9OB3XuQ →
    Details
    Excerpt
    Feizi presented a framework for converting production failures into testable, regression-aware improvements. Key insight: raw session logs plus feedback aren't enough—you need replayable learning environments to make failures actionable.
    Key points
    • Three layers for agent improvement: model weights, harness context, and memory
    • Feizi argues good learning engines seek the smallest durable change at the right layer
    • The core challenge is turning raw session logs into testable replayable environments
    • Trace-to-harness approaches (asking a coding agent to analyze logs) are wipe-based but not testable
    Provenance
    Video · Supporting source
  6. 6

    CuiMao on Seedance 4K for TVC ad production

    X CuiMao

    x.com/CuiMao/status/2073754586775716236 →
    Details
    Engagement
    40 likes · 1 retweets · 16 replies
    Provenance
    Tweet · Primary source