大模型科学代理看似在做研究,实则缺乏科学推理的自我修正能力。
AI scientists produce results without reasoning scientifically

- 通过分解模型与框架贡献,发现基础模型主导性能与行为。
- 68%的推理路径忽略证据,26%未基于反驳修正信念。
- 即使给完整正确推理作为上下文,问题仍持续存在,适合关注可信AI的读者。
基于大语言模型(LLM)的科学代理正被用于自主科研,但其推理是否符合科学探索中使知识可自我修正的认知规范仍不明确。我们评估了跨八个领域的LLM科学代理,涵盖工作流执行到假设驱动探究,共超过25,000次运行,并采用双重分析视角:(i) 系统性性能分析,分解基础模型与代理架构的贡献;(ii) 对代理推理的元认知结构进行行为分析。结果显示,基础模型是性能与行为的主要决定因素,解释方差达41.4%,而架构仅占1.5%。在所有配置中,68%的推理轨迹忽略证据,26%出现基于反驳的信念修正,且多重验证证据极为罕见。该模式在计算工作流与假设探究中均一致存在,即使提供近乎完整的成功推理轨迹作为上下文也未能消除。此类不可靠性在认知要求高的领域中随重复试验累积放大。因此,当前的LLM代理虽能执行科学流程,却未体现科学推理的核心特征。基于结果的评估无法识别这些缺陷,仅靠架构优化亦无法修复。除非将推理过程本身纳入训练目标,否则这些代理产出的科学知识无法由生成过程加以正当化。
原文摘要 · Abstract (English)
Large language model (LLM)-based systems are increasingly deployed to conduct scientific research autonomously, yet whether their reasoning adheres to the epistemic norms that make scientific inquiry self-correcting is poorly understood. Here, we evaluate LLM-based scientific agents across eight domains, spanning workflow execution to hypothesis-driven inquiry, through more than 25,000 agent runs and two complementary lenses: (i) a systematic performance analysis that decomposes the contributions of the base model and the agent scaffold, and (ii) a behavioral analysis of the epistemological structure of agent reasoning. We observe that the base model is the primary determinant of both performance and behavior, accounting for 41.4% of explained variance versus 1.5% for the scaffold. Across all configurations, evidence is ignored in 68% of traces, refutation-driven belief revision occurs in 26%, and convergent multi-test evidence is rare. The same reasoning pattern appears whether the agent executes a computational workflow or conducts hypothesis-driven inquiry. They persist even when agents receive near-complete successful reasoning trajectories as context, and the resulting unreliability compounds across repeated trials in epistemically demanding domains. Thus, current LLM-based agents execute scientific workflows but do not exhibit the epistemic patterns that characterize scientific reasoning. Outcome-based evaluation cannot detect these failures, and scaffold engineering alone cannot repair them. Until reasoning itself becomes a training target, the scientific knowledge produced by such agents cannot be justified by the process that generated it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。