arXiv:2605.30329cs.LG2026-05被引 8

测试AI能否提前识别科研想法的可行性,发现其普遍高估差想法。

SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?

论文配图:SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?
图 1 · 摘自论文原文
  • 构建1099个来自ICLR论文的可复现研究提案数据集
  • 12个前沿大模型普遍高估低质量提案的科学合理性
  • 适合关注AI科研自动化评估可靠性的研究人员

自主AI科研代理旨在通过自动化从假说生成到同行评审的研究流程加速科学发现。然而,现有基准很少检验一个根本瓶颈:大型语言模型能否在耗费时间和计算资源前判断研究想法的方法学可行性。我们提出了SoundnessBench,一个由1099个从ICLR投稿重构的机器学习研究提案组成的精选基准,附有评审员给出的可复现性子评分,并与原始论文交叉验证。SoundnessBench应被视为对可复现提案阶段科学性的评估基准,而非对完整论文评审结果的精确预测。在12个前沿LLM中,我们发现普遍存在乐观偏差:在标准提示下,模型常将低科学性提案误判为合理;而采用激进提示则使错误类型从假阳性转向假阴性。针对公开语料污染、论文标识短语、表面特征及人工审计质量的额外控制表明,该行为并非由单一混淆因素解释。结果表明,当前大模型尚不可作为科学严谨性的独立首道评估者。

原文摘要 · Abstract (English)

Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. However, existing benchmarks rarely test a fundamental bottleneck: whether Large Language Models can judge the methodological viability of a research idea before expending time and computational resources. We introduce SoundnessBench, a curated benchmark of 1,099 machine-learning research proposals reconstructed from ICLR submissions, labeled with reviewer soundness sub-scores, and audited against source papers. SoundnessBench should be interpreted as a benchmark for recoverable proposal-stage soundness rather than exact prediction of full-paper review outcomes. Across 12 frontier LLMs, we find a pervasive optimism bias: under standard prompting, models frequently rate low-soundness proposals as sound, while aggressive prompting largely shifts errors from false positives to false negatives. Additional controls for public-corpus contamination, paper-identifying phrases, surface features, and human audit quality suggest that this behavior is not explained by a single confounder. Our results indicate that current LLMs are not yet reliable as standalone first-gate evaluators for scientific rigor.

AI科研大模型评估科学可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。