Archive BRAID
Agreement Is Not Accuracy / DISPATCH 102
PDF RSS

Dispatch 102 · 2026-07-31 GSV The Instrument Needed Its Own Instrument

Agreement Is Not Accuracy

/ 00:30:18 / 16 sources

“A model can agree with itself, and different models can agree with each other, out of shared bias, a memorized heuristic, or an option-position prior rather than truth.”

— Lenar Kess, today's narration

A revision sweep through the cs.AI batch surfaced five independent papers pushing on the same thing: the instruments we use to decide whether a model is right are themselves unreliable — self-consistency, calibration error, count-based F1, and the model judges grading all of it.

Chapters

  1. 00:00:04 Transcript

Sources

16 cited
  1. 1

    arXiv cs.AI - Research Science (GLOBAL)

    Article

    A new adapter (SAR) that allows models to reliably self-report hidden/misaligned behaviors is a major tool for model auditing and safety, directly impacting how practitioners build and deploy AI.

    arxiv.org/abs/2607.03640 →
    Details
    Context
    A new adapter (SAR) that allows models to reliably self-report hidden/misaligned behaviors is a major tool for model auditing and safety, directly impacting how practitioners build and deploy AI.
    Key points
    • A new adapter (SAR) that allows models to reliably self-report hidden/misaligned behaviors is a major tool for model auditing and safety, directly impacting how practitioners build and deploy AI.
    Provenance
    Article · Supporting source
  2. 2

    arXiv cs.AI - Research Science (GLOBAL)

    Article

    Introduces ANTAP, an evaluation-driven routing architecture for MAS. It addresses critical security vulnerabilities in agent orchestration by replacing text proxies with active capability testing.

    arxiv.org/abs/2606.30555 →
    Details
    Context
    Introduces ANTAP, an evaluation-driven routing architecture for MAS. It addresses critical security vulnerabilities in agent orchestration by replacing text proxies with active capability testing.
    Key points
    • Introduces ANTAP, an evaluation-driven routing architecture for MAS. It addresses critical security vulnerabilities in agent orchestration by replacing text proxies with active capability testing.
    Provenance
    Article · Supporting source
  3. 3

    arXiv cs.AI - Research Science (GLOBAL)

    Article

    Presents a novel hybrid training algorithm (ATOD) for multi-turn agents, directly addressing key limitations in current agentic RL/distillation methods.

    arxiv.org/abs/2606.27814 →
    Details
    Context
    Presents a novel hybrid training algorithm (ATOD) for multi-turn agents, directly addressing key limitations in current agentic RL/distillation methods.
    Key points
    • Presents a novel hybrid training algorithm (ATOD) for multi-turn agents, directly addressing key limitations in current agentic RL/distillation methods.
    Provenance
    Article · Supporting source
  4. 4

    arXiv cs.AI - Research Science (GLOBAL)

    Article

    Presents a novel framework (FCGraft) for code-policy synthesis in embodied agents, directly addressing key limitations (latency/robustness) of current CodeLLMs.

    arxiv.org/abs/2606.13097 →
    Details
    Context
    Presents a novel framework (FCGraft) for code-policy synthesis in embodied agents, directly addressing key limitations (latency/robustness) of current CodeLLMs.
    Key points
    • Presents a novel framework (FCGraft) for code-policy synthesis in embodied agents, directly addressing key limitations (latency/robustness) of current CodeLLMs.
    Provenance
    Article · Supporting source
  5. 5

    arXiv cs.AI - Research Science (GLOBAL)

    Article

    Presents a novel closed-loop training framework (Ouroboros-Spatial) that directly addresses data inefficiency in MLLMs for spatial reasoning. This changes how models are trained and is highly relevant to frontier model…

    arxiv.org/abs/2606.11719 →
    Details
    Context
    Presents a novel closed-loop training framework (Ouroboros-Spatial) that directly addresses data inefficiency in MLLMs for spatial reasoning. This changes how models are trained and is highly relevant to frontier model development.
    Key points
    • Presents a novel closed-loop training framework (Ouroboros-Spatial) that directly addresses data inefficiency in MLLMs for spatial reasoning. This changes how models are trained and is highly relevant to frontier model development.
    Provenance
    Article · Supporting source
  6. 6

    arXiv cs.AI - Research Science (GLOBAL)

    Article

    A unified cross-paradigm benchmark (VendorBench-100) for deepfake detection is a primary artifact that changes how developers evaluate and build safety/detection tools.

    arxiv.org/abs/2607.06254 →
    Details
    Context
    A unified cross-paradigm benchmark (VendorBench-100) for deepfake detection is a primary artifact that changes how developers evaluate and build safety/detection tools.
    Key points
    • A unified cross-paradigm benchmark (VendorBench-100) for deepfake detection is a primary artifact that changes how developers evaluate and build safety/detection tools.
    Provenance
    Article · Supporting source
  7. 7

    arXiv cs.AI - Research Science (GLOBAL)

    Article

    A new, high-performance detection framework (G2VD) for AI video generation is a major artifact with clear security/policy implications.

    arxiv.org/abs/2607.04607 →
    Details
    Context
    A new, high-performance detection framework (G2VD) for AI video generation is a major artifact with clear security/policy implications.
    Key points
    • A new, high-performance detection framework (G2VD) for AI video generation is a major artifact with clear security/policy implications.
    Provenance
    Article · Supporting source
  8. 8

    arXiv cs.AI - Research Science (GLOBAL)

    Article

    Introduces RWGBench, a citation-centric benchmark for LLM related work generation. This addresses a specific, high-signal scholarly writing capability gap.

    arxiv.org/abs/2606.24894 →
    Details
    Context
    Introduces RWGBench, a citation-centric benchmark for LLM related work generation. This addresses a specific, high-signal scholarly writing capability gap.
    Key points
    • Introduces RWGBench, a citation-centric benchmark for LLM related work generation. This addresses a specific, high-signal scholarly writing capability gap.
    Provenance
    Article · Supporting source
  9. 9

    arXiv cs.AI - Research Science (GLOBAL)

    Article

    Addresses a key technical challenge (temporal grounding) for LALMs using a scalable data construction pipeline. Relevant to building advanced multimodal agents.

    arxiv.org/abs/2607.04383 →
    Details
    Context
    Addresses a key technical challenge (temporal grounding) for LALMs using a scalable data construction pipeline. Relevant to building advanced multimodal agents.
    Key points
    • Addresses a key technical challenge (temporal grounding) for LALMs using a scalable data construction pipeline. Relevant to building advanced multimodal agents.
    Provenance
    Article · Supporting source
  10. 10

    arXiv cs.AI - Research Science (GLOBAL)

    Article

    Addresses the core 'alignment target problem' by showing human moral judgment diverges based on AI origin (human vs. machine vs. designer). This is a key debate in safety/governance.

    arxiv.org/abs/2604.24155 →
    Details
    Context
    Addresses the core 'alignment target problem' by showing human moral judgment diverges based on AI origin (human vs. machine vs. designer). This is a key debate in safety/governance.
    Key points
    • Addresses the core 'alignment target problem' by showing human moral judgment diverges based on AI origin (human vs. machine vs. designer). This is a key debate in safety/governance.
    Provenance
    Article · Supporting source
  11. 11

    Japan Digital Agency News - Policy Geopolitics (JP)

    Article

    Updates on digital identity/MyNumber systems are key policy areas in AI infrastructure and governance, showing state adoption of digital records.

    www.digital.go.jp/policies/mynumber/mynumbe… →
    Details
    Context
    Updates on digital identity/MyNumber systems are key policy areas in AI infrastructure and governance, showing state adoption of digital records.
    Key points
    • Updates on digital identity/MyNumber systems are key policy areas in AI infrastructure and governance, showing state adoption of digital records.
    Provenance
    Article · Supporting source
  12. 12

    When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

    Source Kaihua Ding

    Agreement is not accuracy: a model can agree with itself, and different models can agree with each other, out of shared bias, a memorized heuristic, or an option-position prior rather than truth.

    arxiv.org/abs/2607.08065 →
    Details
    Cited text
    Agreement is not accuracy: a model can agree with itself, and different models can agree with each other, out of shared bias, a memorized heuristic, or an option-position prior rather than truth.
    Key points
    • 53 runners, K=50 samples per case, GPQA Diamond and AIME, 265,000 samples total.
    • Agreement is a positive but weak predictor of majority-correctness, rho 0.20 to 0.59, positive under item-clustered resampling.
    • Worst regime is the most consistent frontier model: agreement at or above 0.8 on 77% of GPQA case-result entries, 48% of those wrong.
    • Version history from the abstract page: v1 9 Jul 2026, v2 28 Jul 2026 — this is a revision, not a new publication.
    Provenance
    Source · Background source
  13. 13

    Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory

    Source Jon-Paul Cacioli

    This version (v3) corrects a differential length bias in the automated correctness scorer, validated against 1,830 human adjudications; the inverse accuracy-efficiency coupling reported in v1 and v2 does not survive rel…

    arxiv.org/abs/2603.25112 →
    Details
    Cited text
    This version (v3) corrects a differential length bias in the automated correctness scorer, validated against 1,830 human adjudications; the inverse accuracy-efficiency coupling reported in v1 and v2 does not survive relabelling.
    Key points
    • 4 models, 224,000 factual QA trials; normalised metacognitive information used in place of meta-d prime because open-ended QA has no two-alternative Type-1 decision.
    • Metacognitive information varies by a factor of 1.98; rank correlation with accuracy is -0.80 on TriviaQA and +0.00 on Natural Questions.
    • Efficiency is weakest in Science and Technology for every model tested; meta-I tracks abstention gain at rho +1.00 while accuracy does not.
    • The v3 self-correction was not in the packet — found on the abstract page version note.
    Provenance
    Source · Background source
  14. 14

    Prompt Framing Distorts Count Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

    Source Dekun Yang

    Count-based F1 can rise dramatically without any improvement in span localization, a phenomenon we term F1 Inflation.

    arxiv.org/abs/2607.01240 →
    Details
    Cited text
    Count-based F1 can rise dramatically without any improvement in span localization, a phenomenon we term F1 Inflation.
    Key points
    • ErrorBench: 6 models, 5 prompt conditions, 4,290 responses over 143 CoNLL-2014 passages.
    • Anchored prompts produce up to 0.79 points of F1 inflation under M2 scoring, up to 0.96 under strict matching.
    • ERRANT 3.0.0 replication on 100 passages: Blind to Anchored raises Count-F1 by +0.21 but multi-reference ERRANT F0.5 by only +0.04.
    • Larger count responses come from the highly instruction-compliant GPT and Claude systems; smaller from the Gemini family.
    Provenance
    Source · Background source
  15. 15

    Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control

    Source Brett Reynolds

    The intended use is diagnosis, not deployment certification, vendor ranking, or a general safety score.

    arxiv.org/abs/2607.01153 →
    Details
    Cited text
    The intended use is diagnosis, not deployment certification, vendor ranking, or a general safety score.
    Key points
    • 18-item seed benchmark, 54-row pilot, six-cell LLM-judge assessment.
    • The first judge graded its own outputs with the expected answer visible and missed the safety-relevant minority classes.
    • Rejudging across 3 judge models and 2 information conditions: no cell recovers more than two of eleven partial successes; the strongest cell's edge comes partly from never using that label.
    • Item-clustered intervals leave four of six chance-corrected agreement statistics unable to rule out a constant labeller.
    Provenance
    Source · Background source
  16. 16

    FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

    Source Yufeng Wang

    This paper has been withdrawn by Yufeng Wang

    arxiv.org/abs/2607.05682 →
    Details
    Cited text
    This paper has been withdrawn by Yufeng Wang
    Key points
    • Withdrawal confirmed on the abstract page: v2, 28 Jul 2026, comment reads "Withdraw for further improvement and work consolidation"; page states no license for this version due to withdrawn.
    • The withdrawal was not flagged in the day's candidate pool — found by fetching the abstract page directly.
    • Core artifact is a Research Question Certificate recording primitives, assumptions, a mechanism model, a tension, a falsifiable hypothesis, a minimal decisive test, and a failure update rule.
    • Scored 4.86/5 versus 4.38/5 for the strongest baseline under a DeepSeek blind judge, with a Gemini-2.5-Flash rescore at Pearson 0.865; removing certificates drops below 1/5.
    Provenance
    Source · Background source