让AI科研过程可追溯,防止结论漂移
Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

- 将文献、思路、实验等转化为可审查的持久化科研成果
- 在无训练记忆系统中保持从问题到验证的可追踪路径
- 适合关注AI科研可信性与可复现性的研究者
AI系统正越来越多地自动化科学工作流程,但连接已有证据、生成想法、实验与最终结论之间的推理过程仍常隐含于模型推断中。本文提出Xcientist——一种研究框架,将研究综合与实验验证外化为可审查、受合约约束的过程。Xcientist将文献证据、想法状态、实施计划、消融记录和修复痕迹组织为持久的研究成果,使生成机制可在不丢失证据基础的前提下被验证、执行与修订。我们识别出‘结论漂移’作为自动化研究的失败模式:可运行的成果不再支持最初声称的机制。在无需训练的记忆系统、基于图结构的交通预测以及多尺度物理信息神经网络中,Xcientist均实现了从问题提出到机制设计、验证与有限修订的可追踪轨迹。结果表明,评估AI科学家不仅要看最终成果,更应考察其综合与验证过程是否可归因、可审查且具有科学问责性。
原文摘要 · Abstract (English)
AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside model inference. Here we introduce Xcientist, a research harness that externalizes research synthesis and experimental validation into inspectable, contract-governed processes. Xcientist organizes literature evidence, idea states, implementation plans, ablation records and repair traces as persistent research artifacts, so that generated mechanisms can be grounded, executed, tested and revised without losing their evidential basis. We identify claim drift as a failure mode of automated research, where runnable artifacts no longer support the mechanism originally claimed. Across training-free memory systems, graph-structured traffic forecasting and multi-scale physics-informed neural networks, Xcientist preserves traceable trajectories from problem formulation to mechanism design, validation and bounded revision. These results suggest that AI scientists should be evaluated not only by their final artifacts, but by whether their synthesis and validation processes remain attributable, inspectable and scientifically accountable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。