arXiv:2604.22080cs.AI2026-04

用智能体主动找反例,让科学结论更可靠

Sound Agentic Science Requires Adversarial Experiments

  • 让智能体不只生成论证,更主动寻找推翻假设的证据
  • 揭示当前研究中因选择性分析导致的虚假可信结果
  • 适合追求严谨科学推理的研究者和审稿人

基于大模型的智能体正被广泛用于科学数据分析,自动化原本依赖人力和专业知识的任务。这一能力虽被视为加速发现,却也加剧了常见问题:快速生成看似合理、可无限修改的分析,仅选取支持发表结果的数据,将假设空间变为可被选择性验证的假说。与软件不同,科学知识不能仅靠代码迭代和事后统计验证。流畅的解释或单一数据集上的显著结果并非验证。缺失的证据是未被检验的否定性空间——那些本可能证伪主张的实验从未执行或未发表。因此我们提出:对由智能体生成的非实验性主张,应采用‘先证伪’标准:智能体不应主要用于构建最具说服力的叙事,而应主动探索主张失败的各种方式。

原文摘要 · Abstract (English)

LLM-based agents are rapidly being adopted for scientific data analysis, automating tasks once limited by human time and expertise. This capability is often framed as an acceleration of discovery, but it also accelerates a familiar failure mode, the rapid production of plausible, endlessly revisable analyses that are easy to generate, effectively turning hypothesis space into candidate claims supported by selectively chosen analyses, optimized for publishable positives. Unlike software, scientific knowledge is not validated by the iterative accumulation of code and post hoc statistical support. A fluent explanation or a significant result on a single dataset is not verification. Because the missing evidence is a negative space, experiments and analyses that would have falsified the claim were never run or never published. We therefore propose that non-experimental claims produced with agentic assistance be evaluated under a falsification-first standard: agents should not be used primarily to craft the most compelling narrative, but to actively search for the ways in which the claim can fail.

智能体科学验证因果推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。