arXiv:2609.07611cs.AIcs.CL2026-09

新基准评估AI在主动探索中提出科学假说的能力,发现强模型提升更明显。

AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era

  • 设计主动探索与静态观察双场景对比评测
  • 强模型在主动探索中能力提升速度是静态的两倍
  • 适合研究AI科学家推理能力及能力差距的团队

科学构想能力是基于科学证据生成新颖可验证假说的能力,对自主AI科学家至关重要。现有评估多依赖静态参考文献生成想法,脱离现代AI科学家的检索-推理流程,且随模型进步变得不敏感。我们提出AgentIdeaBench,一个跨五学科、40个细分领域的多维度基准,包含静态观察与主动探索两种匹配设置。对33个大模型进行评估,采用文献验证的评分框架,由评审者对比已有研究评估原创性。主动探索揭示更大能力空间,且该空间分布不均:性能提升约快两倍,且收益受模型能力制约,强者获益更多。探索提升主要来自更好的立足点,改善了可行性、清晰度和具体性,但原创性评分不变。我们还引入科学世界建模——生成时通过结构化思想实验迭代优化假设,对中等能力模型有效,而前沿模型已内化此类推理,效果减弱。AgentIdeaBench为智能体时代的科学构想研究提供了适配的评估基础。

原文摘要 · Abstract (English)

Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely assess it by asking models to generate ideas from a static, curated set of reference papers. That passive setup departs from the retrieval-and-reasoning workflow of modern AI scientists, and it becomes less discriminative as models improve. We introduce AgentIdeaBench, a multidisciplinary benchmark that evaluates scientific ideation under two matched settings, static observation and active exploration. We report matched Static-Active evaluations for 33 LLMs across 40 densely scored subfields spanning five disciplines, using a multidimensional, literature-verified scoring framework whose critics assess originality against retrieved prior art. Active exploration reveals considerably more capability headroom, and that headroom is unevenly distributed across models. Performance scales about twice as fast as under static observation, and the exploration gain is capability-gated, favoring the strongest models over the weakest. The gain reflects better grounding, improving feasibility, clarity, and specificity while leaving measured originality unchanged under our critics. We further explore Scientific World Modeling, a generation-time loop that refines a draft hypothesis through structured thought experiments. It benefits mid-capability models, and its impact diminishes among frontier models that appear to have internalized such reasoning patterns already. AgentIdeaBench gives future work on scientific ideation a measurement basis suited to the agent era.

科学构想大模型评估主动探索智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。