现有智能体无法真正自主科研,因存在认知、知识与机制缺陷。
Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

- 智能体依赖大模型,但训练数据缺失实验失败经验
- 偏好优化使输出趋同,抑制科学探索多样性
- 适合关注科研自动化局限的学者与系统设计者
越来越多研究致力于构建端到端自主科研的AI科学家。本文指出,尽管它们已可作为科研协作者,但当前架构并不具备真正自主发现的能力。主要挑战包括:(1)问题选择受麦纳马拉谬误影响;(2)基于大语言模型的智能体缺乏实验室操作中的隐性程序知识与失败经验;(3)后训练阶段的偏好优化压缩输出多样性,趋向共识;(4)多数科学基准仅评估单轮预测准确率,缺乏物理实验对计算模型的反馈。这些并非单纯规模或结构问题,需重新审视基础设计。建议采用科学模拟作为训练验证器,构建能反映科研目标演化的持久世界模型,建立集中式预注册库以记录所有AI生成假设,并让应用由科学需求驱动而非工具便利性。
原文摘要 · Abstract (English)
A growing body of work pursues AI scientists capable of end-to-end autonomous scientific discovery. This position paper argues that although they already function as co-scientists, agentic AI scientists are not built for autonomous scientific discovery. We identify the following challenges in building and deploying autonomous AI scientists: (1) Problem selection is influenced by the McNamara fallacy; (2) Agents are built on large language models (LLMs) whose training corpora omit tacit procedural and failure knowledge of laboratory practice; (3) Preference optimisation during post-training compresses output diversity toward consensus; and (4) Most scientific benchmarks measure single-turn prediction accuracy and lack feedback from physical experiments back to the computational model. These challenges are not just questions of scale and scaffolding; they require revisiting fundamental design choices. To build truly autonomous AI scientists, we recommend the use of scientific simulations as verifiers for training, the design of persistent world models that represent the shifting objectives governing real investigations, the establishment of a centralized preregistration repository for all AI-generated hypotheses, and application driven by scientific need rather than tool affordance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。