AI可自动复现并发现物理论文方法问题,96.6%的批评来自实际计算而非阅读。
Grounded autonomous scrutiny at scale: emergent critique from reproduction of published computational physics papers

- AI读论文+重现实验,从执行中自动发现方法缺陷
- 111篇论文中42%出现实质性问题,96.6%需运行计算才暴露
- 能生成完整评论,挑战原论文关键结论,适合审稿与科研验证
自主大模型代理可在机器学习沙盒中生成完整研究成果,但真实计算物理更复杂:实验基于第一性原理计算,依赖可重演的物理真值,且新工作几乎总建立在关键已有论文之上。我们探究此类代理能否对已发表的计算物理论文进行有根基的审视——即阅读论文、从零复现,并在执行中揭示方法论问题。部署单一Claude Opus 4.6配置于两个互补层面:在规模上,覆盖111篇开源的Quantum ESPRESSO论文,代理自动执行读-规划-计算-比对循环,虽未被要求批判,仍对约42%的论文提出实质性方法论质疑;其中88篇中的85篇(96.6%)仅在实际运行计算后浮现,阅读阶段仅识别出1.8%。批判源于复现,而非阅读。在深度层面,针对一篇《Nature Communications》关于二维材料MOSFET多尺度器件模拟的论文,一个继承已验证复现流程的新代理自动生成包含14个问题的物理清单,并撰写一份完整的六页投稿格式评论,修正了原文中L_G = 5 nm的核心结论。其两项挑战该结论的攻击——源退化接触电阻上限与锑掺杂退化率——在原始21位审稿人评审中均未出现。
原文摘要 · Abstract (English)
Autonomous LLM agents now produce complete research artifacts in machine-learning sandboxes, but real computational physics is harder: experiments are first-principles calculations against re-runnable physical ground truth, and meaningful new work almost always builds on a key existing paper. We ask whether such an agent can perform grounded scrutiny of published computational physics - reading a paper, reproducing it from scratch, and surfacing methodological concerns from execution. We deploy a single Claude Opus 4.6 configuration at two complementary scopes. At scale, across 111 open-access Quantum ESPRESSO papers, an autonomous agent runs the read-plan-compute-compare loop and, although never asked to critique, raises substantive methodological concerns on ~42% of papers; 85 of 88 of these critiques (96.6%) surface only after the agent has actually run a calculation, with a reading-only ceiling of 1.8%. Critique emerges from reproduction, not from reading. In depth, on one Nature Communications paper on multiscale device simulation of a 2D-material MOSFET, a fresh agent inheriting a verified reproduction pipeline autonomously produces a 14-concern physics inventory and a complete, submission-form six-page Comment that revises the paper's L_G = 5 nm headline. Two of its L_G = 5 nm headline-challenging attacks - a source-degeneration contact-resistance bound and a Sb-doping degradation ratio - are absent from the published 21-reviewer peer review.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。