arXiv:2606.31478cs.AIcs.CV2026-06被引 1

让科研代理通过多假设诊断失败,提升实验纠错能力。

One Reflection Is Not Enough: Self-Correcting Autonomous Research via Multi-Hypothesis Failure Attribution

论文配图:One Reflection Is Not Enough: Self-Correcting Autonomous Research via Multi-Hypothesis Failure Attribution
图 1 · 摘自论文原文
  • 用多假设分析故障原因,分层定位问题根源。
  • 生成结果准确率从42%提升至92%,论文质量达6.75/10。
  • 适合追求高可靠性的自动化科研系统研究者。

自主科研代理虽能提出假说、编写代码、运行实验并产出论文,但实验失败时仍易出错。当前主流方法依赖单一自由形式反思,将复杂日志压缩为一段文字,常导致局部试错或抛弃有用信息的硬转向。本文提出SAGE(自校正、自主、基于证据的实验者),核心机制为多假设故障归因(MHFA),将恢复过程转化为结构化因果诊断。通过分析动态轨迹特征,MHFA系统性生成多个有证据支持的失败解释,独立评估其严重性,并确定性地将根因导向正确干预层级(假说、实验设计或实现)。为保障科学诚实,SAGE采用基于证据的报告机制,强制限制生成结果仅限实际测量值,删去虚构数值。在12个主题、5个领域的基准测试中,相较于反思基线,SAGE将可承载指标的输出从42%提升至92%,成果质量从5.00提升至6.75/10,盲测超越AI-Scientist-v2(52.0 vs. 48.2),优势集中在代码开发与执行环节。尽管完全自主撰写会议级论文仍是开放难题,SAGE仍显著提升了科研成果的可靠性与质量。通过结构化恢复与显式约束结合,SAGE大幅优于传统单一体反思范式,为未来自主科研奠定了高可信基础。

原文摘要 · Abstract (English)

Autonomous research agents can now draft hypotheses, write code, run experiments, and produce papers, but they remain brittle when experiments fail. Under the prevailing paradigm, failure recovery is usually delegated to a single free-form reflection: a rich trajectory of metrics, logs, and design choices is compressed into one verbal critique, which often leads either to localized trial-and-error or to hard pivots that discard useful context. We propose SAGE, a Self-correcting, Autonomous, Grounded Experimenter, to tackle this failure-recovery bottleneck. Its core mechanism, Multi-Hypothesis Failure Attribution (MHFA), treats recovery as a structured causal diagnosis. By analyzing dynamic trajectory features, MHFA systematically generates multiple evidence-grounded explanations for a failure, independently evaluates their severity, and deterministically routes the verified root cause to the correct intervention level (hypothesis, experimental design, or implementation). To guarantee scientific honesty, SAGE further employs a grounded reporting mechanism that explicitly constrains drafted results to actual measured values, redacting hallucinated numbers. On a 12-topic, 5-domain benchmark, SAGE increases metrics-bearing outputs from 42% to 92% over a reflection baseline, improves artifact quality from 5.00 to 6.75/10, and blindly outscores AI-Scientist-v2 (52.0 vs. 48.2), with gains concentrated in code development and execution. While fully autonomous scientific writing and generating conference-ready papers remain notoriously difficult open problems for the entire field, SAGE successfully produces significantly more reliable and higher-quality scientific artifacts. Ultimately, by coupling structured recovery with explicit grounding constraints, SAGE significantly outperforms monolithic reflection paradigms, establishing a highly trustworthy foundation for future autonomous research.

自主科研故障诊断自动化实验可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。