提出可验证的自主科研系统,让机器写作不再造假。
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

- 用证据链确保每个结论都有据可查,从文献到代码全程追踪。
- 在75篇论文测试中,零虚构引用、100%得分可验证、代码方法对齐率最高。
- 适用于需要高可信度科研自动化场景,如医学影像与语言建模。
自主研究代理虽能产出媲美人类的成果和论文,但其内容常存在不可靠问题:虚构引用、无法复现的性能指标、方法描述与代码不一致。为此,本文提出三项贡献:第一,链式证据(Chain-of-Evidence, CoE)框架,要求每个主张都可追溯至原始证据;第二,ScientistOne系统,一个端到端的自主研究系统,在文献综述、方案发现和论文撰写中自始至终维护证据链;第三,CoE审计,一种后置审计机制,包含得分验证、规范违背检测、参考文献验证和方法-代码一致性检查,统一应用于所有系统。在涵盖五个前沿任务的75篇论文测试中,所有基线系统均出现至少一种系统性失败:虚构引用率达21%,得分验证通过率低至42%,方法-代码对齐率介于20%至80%之间。ScientistOne实现零虚构引用(0/337)、100%得分验证通过(12/12),并取得最高方法-代码对齐率(14/15),在五项任务上达到或超越人类专家水平。该系统还拓展至六项额外任务,覆盖医学影像、细粒度识别、3D感知与语言建模,于Parameter Golf达最先进水平,在MLE-Bench任务中获得金牌,而基线系统则完全失败。
原文摘要 · Abstract (English)
Autonomous research agents produce competitive solutions and professional-looking manuscripts, yet their outputs contain verifiability failures undetectable by surface-level evaluation: fabricated citations, unreproducible scores, and method descriptions that diverge from the implementation. We address this through three contributions. First, Chain-of-Evidence (CoE), a verifiability framework requiring every claim to be traceable to its evidence source. Second, ScientistOne, an end-to-end autonomous research system that maintains evidence chains by construction throughout literature review, solution discovery, and paper writing. Third, CoE Audit, a post-hoc audit whose four integrity checks -- score verification, specification violation, reference verification, and method-code alignment -- apply uniformly to all systems. Across 75 papers spanning five systems and five frontier research tasks, every baseline exhibits at least one systematic failure mode: hallucinated reference rates reach 21%, score verification passes in as few as 42% of papers, and method-code alignment ranges from 20% to 80%. ScientistOne achieves zero hallucinated references (0/337), perfect score verification (12/12), and the highest method-code alignment (14/15), while matching or exceeding human expert performance on all five tasks. ScientistOne further generalizes to six additional tasks spanning medical imaging, fine-grained recognition, 3D perception, and language modeling, achieving state-of-the-art on Parameter Golf and gold medals on MLE-Bench tasks where baselines fail entirely.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。