提出真实多物理场流的神经代理模型评测框架,揭示现有模型在复杂场景下的脆弱性
Benchmarking neural surrogates on realistic spatiotemporal multiphysics flows
- 构建涵盖11个高保真数据集的REALM评测框架,覆盖从基础问题到推进与防火场景
- 发现模型性能受限于维度、刚度和网格不规则性,错误随规模迅速增长
- 模型准确率高但常遗漏关键瞬态结构,适合关注物理可信性的研究者参考
预测多物理场动力学因多尺度、异质物理过程严重耦合而计算成本高昂且挑战重重。尽管神经代理模型有望带来范式变革,但该领域目前存在‘掌握假象’——顶级评论反复指出,现有评估过度依赖简化、低维代理,无法暴露模型在真实场景中的固有脆弱性。为弥合这一关键差距,我们提出REALM(REalistic AI Learning for Multiphysics)评测框架,用于在具有应用导向的复杂反应流中测试神经代理模型。REALM包含11个高保真数据集,涵盖经典多物理场问题至复杂推进与防火安全场景,并提供标准化端到端训练与评估流程,包含多物理场感知预处理和稳健滚动策略。基于此框架,我们系统评估了十余种代表性代理模型家族,包括谱算子、卷积模型、Transformer、点对点算子以及图/网格网络,识别出三大稳健趋势:(i) 尺度、刚度与网格不规则性共同决定的扩展障碍,导致滚动误差迅速上升;(ii) 性能主要由架构归纳偏置控制,而非参数量;(iii) 名义准确率与物理可信行为间存在持续差距,即使相关性高,模型仍可能遗漏关键瞬态结构与积分量。总体而言,REALM揭示了当前神经代理模型在真实多物理场流中的局限性,并提供了一个严谨的测试平台,推动下一代物理感知架构的发展。
原文摘要 · Abstract (English)
Predicting multiphysics dynamics is computationally expensive and challenging due to the severe coupling of multi-scale, heterogeneous physical processes. While neural surrogates promise a paradigm shift, the field currently suffers from an "illusion of mastery", as repeatedly emphasized in top-tier commentaries: existing evaluations overly rely on simplified, low-dimensional proxies, which fail to expose the models' inherent fragility in realistic regimes. To bridge this critical gap, we present REALM (REalistic AI Learning for Multiphysics), a rigorous benchmarking framework designed to test neural surrogates on challenging, application-driven reactive flows. REALM features 11 high-fidelity datasets spanning from canonical multiphysics problems to complex propulsion and fire safety scenarios, alongside a standardized end-to-end training and evaluation protocol that incorporates multiphysics-aware preprocessing and a robust rollout strategy. Using this framework, we systematically benchmark over a dozen representative surrogate model families, including spectral operators, convolutional models, Transformers, pointwise operators, and graph/mesh networks, and identify three robust trends: (i) a scaling barrier governed jointly by dimensionality, stiffness, and mesh irregularity, leading to rapidly growing rollout errors; (ii) performance primarily controlled by architectural inductive biases rather than parameter count; and (iii) a persistent gap between nominal accuracy metrics and physically trustworthy behavior, where models with high correlations still miss key transient structures and integral quantities. Taken together, REALM exposes the limits of current neural surrogates on realistic multiphysics flows and offers a rigorous testbed to drive the development of next-generation physics-aware architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。