测试物理大模型在多种场景下的泛化能力,发现其表现受条件制约。
Do Physics Foundation Models Learn Generalizable Physics? A Bias-Aware Benchmark Across Physical Regimes and Distribution Shifts

- 构建8种物理动态、25种测试场景的基准,覆盖不同尺度与初始条件
- 6万次测量显示模型泛化能力依赖具体物理场景和预训练条件
- 当前模型仅为条件泛化,需新机制突破跨域迁移瓶颈
近期物理基础模型声称具备通用时空预测能力,但其评估常将性能简化为单一平均分,难以判断是否真正掌握可迁移的物理规律。本文构建包含8种物理动态、3种训练数据混合方式、25种由动态尺度和初始条件复杂度变化引发的测试场景的基准,涵盖分布内、分布偏移和分布外设置。评估五种模型架构及每种架构的四个变体(从头训练与三种预训练规模),共产生6万次测量结果。结果显示,当前物理基础模型表现为条件泛化而非通用泛化:其性能受物理场景、时间尺度、初始条件、预训练、模型规模和架构影响显著。仅扩大训练数据分布无法有效缓解此局限,预训练与扩展也未能可靠消除其能力偏差。我们主张,提升物理基础模型需超越单纯扩大模型或数据,转向学习能跨场景、跨尺度、跨分布迁移的可转移物理知识机制。
原文摘要 · Abstract (English)
Recent physics foundation models claim general spatiotemporal forecasting ability, yet their evaluations often collapse performance into a single average score under a fixed training distribution. This makes it difficult to determine whether a model has learned generalizable physical dynamics or only performs well under particular settings. We construct a benchmark with 8 physical dynamics, 3 training-data mixtures, and 25 test regimes induced by dynamic-scale and initial-condition complexity shifts, covering in-distribution, distribution-shift, and out-of-distribution settings. We evaluate five physics foundation model architectures and four model variants per architecture (scratch and three pretrained sizes), resulting in 60,000 measurements. Our results show that current physics foundation models behave as conditional rather than universal generalists: their generality depends on the physical regime, temporal scale, initial-condition setting, pretraining, model size, and architecture. Improving the training data distribution only partially mitigates this limitation. Pretraining and scaling are also unable to reliably remove their ability biases. We argue that improving physics foundation models requires moving beyond scaling models or expanding data, toward learning mechanisms that better capture transferable physical knowledge across regimes, temporal scales, and distribution shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。