arXiv:2509.08713cs.AIcs.DL2025-09被引 20

AI科研系统自动化越强,隐藏缺陷越难发现,需公开全流程日志确保可信。

The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems

  • 通过控制实验识别四大故障模式:基准选择不当、数据泄露等
  • 实测两款主流AI科研系统均存在多类缺陷,严重程度不一
  • 仅看论文难发现问题,必须提交完整代码与日志才可有效检测

AI科研系统能够自主完成从假设生成、实验到论文撰写的完整研究流程,具有加速科学发现的巨大潜力。然而,这些系统内部工作流尚未得到充分审视,可能引入影响研究成果可靠性与可信度的缺陷。本文识别出当代AI科研系统的四种潜在失效模式:不恰当的基准选择、数据泄露、指标误用以及事后选择偏差。为评估这些风险,我们设计了隔离每种失效模式的受控实验,并克服了评估AI科研系统特有的挑战。对两款主流开源AI科研系统的评估显示,存在多种可被忽略的缺陷,严重程度各异。最后,我们证明,获取完整的自动化流程日志和代码,比仅审查最终论文能更有效地发现这些问题。因此,我们建议期刊和会议在评估AI生成的研究时,强制要求提交这些原始资料,以保障透明性、可问责性和可复现性。

原文摘要 · Abstract (English)

AI scientist systems, capable of autonomously executing the full research workflow from hypothesis generation and experimentation to paper writing, hold significant potential for accelerating scientific discovery. However, the internal workflow of these systems have not been closely examined. This lack of scrutiny poses a risk of introducing flaws that could undermine the integrity, reliability, and trustworthiness of their research outputs. In this paper, we identify four potential failure modes in contemporary AI scientist systems: inappropriate benchmark selection, data leakage, metric misuse, and post-hoc selection bias. To examine these risks, we design controlled experiments that isolate each failure mode while addressing challenges unique to evaluating AI scientist systems. Our assessment of two prominent open-source AI scientist systems reveals the presence of several failures, across a spectrum of severity, which can be easily overlooked in practice. Finally, we demonstrate that access to trace logs and code from the full automated workflow enables far more effective detection of such failures than examining the final paper alone. We thus recommend journals and conferences evaluating AI-generated research to mandate submission of these artifacts alongside the paper to ensure transparency, accountability, and reproducibility.

AI科研自动化可信度可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。