arXiv:2607.02329cs.AIcond-mat.mtrl-sci2026-07中稿 · ICML被引 2

用1.1万篇论文训练智能体,自动完成前沿物理研究并生成可发表论文。

Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics

  • 从海量论文中自主发现研究方向,通过复现文献实现方法校准。
  • 完成3项新物性发现,共进行2162次文献查阅,仅在复现失败时需少量人工干预。
  • 通过多阶段容错设计确保结果可信,适合高风险科学领域自动化研究。

自主研究代理在机器学习沙箱中已实现端到端自动化执行,但前沿物理科学存在本质差异:物理推理贯穿方法选择,工具链常缺乏文档,且校准必须依赖外部文献锚点。未受约束的代理虽引用文献却无法验证,易产生基于内部先验的合理但不可验证结果。本文提出一个从11,083篇近期凝聚态物理arXiv论文出发,自动生成出版级论文的全流程管道。该代理自主规划研究方向,通过复现已有文献校准方法,开展首次原理计算,并撰写论文,全程依托文献支撑。整个流程在六个阶段、47个独立上下文会话中完成,共享磁盘状态,共触发2,162次文献咨询。故障容错来自冗余机制:上下文隔离、分布式文献校准与对抗性审查可捕捉单一会话遗漏问题。预演和后演阶段完全自主,仅在复现失败时需有限人工介入——为操作知识整理而非科学方向指导。通过对比基线与无引导消融实验,确认结构化数值验证是校准环节的关键机制。该工作为高风险科学领域的自主研究提供了可量化、可复制的基础框架。

原文摘要 · Abstract (English)

Autonomous-research agents have demonstrated end-to-end LLM automation in machine-learning sandboxes where execution provides calibration. Frontier physical science differs categorically: physical reasoning underlies every methodology choice, toolchains are often underdocumented, and calibration must come from external literature anchors - which unscaffolded agents cite but do not confront, hallucinating plausible, unverifiable results from internal priors. We present a pipeline that runs end-to-end from a corpus of 11,083 recent condensed-matter physics arXiv papers to a publication-grade manuscript with three substantive physics findings (here on altermagnetic piezomagnetism): the agent autonomously conceives a research direction by mapping the corpus, calibrates methodology by reproducing published references, conducts novel first-principles computations, and writes the manuscript - grounded in literature throughout, across 47 fresh-context sessions in six phases sharing only on-disk state, with 2,162 literature-consultation events. Fault tolerance emerges from redundancy: fresh-context isolation, distributed grounding, and adversarial review catch what any single session misses; pre- and post-pilot stages are fully autonomous, and pilot requires bounded human intervention only at reproduction failures - operational knowledge curation, not scientific direction. Two paired failure modes - a pre-architecture baseline and a no-pilot ablation - isolate structurally enforced numerical confrontation at calibration checkpoints as the operative grounding mechanism. The primitives, characterized failure modes, and quantified intervention pattern lay a foundation for autonomous research in high-stakes scientific domains beyond computational physics.

自主研究物理计算大模型容错机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。