arXiv:2602.18458cs.CYcs.AI2026-02被引 4

用AI自动评估论文实验可复现性,发现人类忽略的51个问题

The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research

  • 构建基于代码与数据的自动化评估框架,检验研究真实性
  • 与人类评委一致率达80%以上,识别出51处人工遗漏缺陷
  • 适合关注科研可信度、可复现性的研究人员参考

科学界普遍存在的可复现危机凸显了以论文为中心的评审体系在评估研究严谨性和可复现性方面的局限性。人工智能代理自主设计并生成大量研究产出,进一步加剧了这一挑战。本文提出首个基于执行过程的评估框架,通过检查代码和数据,超越传统叙事式评审。以机制可解释性研究为测试场景,构建标准化研究输出,开发了MechEvalAgent这一自动化评估系统,用于评估实验流程一致性、结果可复现性及结论泛化能力。结果显示,该框架与人类评审者达成超过80%的一致性,识别出显著的方法学问题,并揭示了51项人类评审者未能发现的问题。本工作展示了人工智能代理在变革科研评价中的潜力,为实现更严谨的科学实践铺平道路。

原文摘要 · Abstract (English)

Reproducibility crises across sciences highlight the limitations of the paper-centric review system in assessing the rigor and reproducibility of research. AI agents that autonomously design and generate large volumes of research outputs exacerbate these challenges. In this work, we address the growing challenges of scalability and rigor by flipping the dynamic and developing AI agents as research evaluators. We propose the first execution-grounded evaluation framework that verifies research beyond narrative review by examining code and data alongside the paper. We use mechanistic interpretability research as a testbed, build standardized research output, and develop MechEvalAgent, an automated evaluation framework that assesses the coherence of the experimental process, the reproducibility of results, and the generalizability of findings. We show that our framework achieves above 80% agreement with human judges, identifies substantial methodological problems, and surfaces 51 additional issues that human reviewers miss. Our work demonstrates the potential of AI agents to transform research evaluation and pave the way for rigorous scientific practices.

可复现性AI评估科研可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。