AI解偏微分方程时,准确不等于有科学证据,该论文首次评估模型输出能否支撑科学结论。
When Does AI for PDEs Yield Scientific Evidence?

- 提出新评估框架,检验AI模型输出是否能支持具体科学命题
- 发现数值精度高的模型未必提供更强科学证据,两者排名常不一致
- 揭示现有基准可能偏向证据支持力弱的方法,适合关注科学可信度的研究者
现有AI解偏微分方程(PDE)的基准主要衡量预测或逼近精度。但在物理研究中,AI输出常需作为科学主张的证据。这两者并不等同:前者衡量输出与参考目标的一致性或约束满足程度;后者则判断在给定研究对象、科学主张、假设和证据标准下,输出是否足以支持该主张。为弥合这一差距,我们扩展了广泛使用的PDE仿真基准和全面的PDE反问题基准,首次在AI for PDE领域实现对模型输出能否支持特定科学主张的评估。结果表明,数值精度与证据支持力可导致模型排名不同,我们解释了何时何地出现这种差异,并发现现有基准可能偏好那些对科学主张支持力较弱的方法。综上,我们形式化、实证并解释了AI for PDE中的评估-用途错配问题。
原文摘要 · Abstract (English)
Existing AI-for-PDE benchmarks primarily assess models in terms of predictive or approximation accuracy. In physics research, however, AI outputs often serve as evidence for scientific claims. These two objectives are not equivalent: the former measures an output's agreement with a reference target or satisfaction of governing constraints; the latter asks whether, given a specified object of study, scientific claim, assumptions, and evidence standard, the output provides sufficient evidence for that claim. To bridge this gap, we extend a widely used PDE-simulation benchmark and a comprehensive benchmark for PDE inverse problems to enable, for the first time in AI for PDEs, evaluation of whether and to what extent model outputs support specified scientific claims. Our results show that numerical accuracy and evidential support can rank models differently, explain when and why they do so, and reveal that existing benchmarks can favor methods whose outputs provide weaker support for the scientific claims of interest. Together, we formalize, empirically demonstrate, and explain this evaluation--use mismatch in AI for PDEs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。