arXiv:2509.00024physics.comp-phcs.LG2025-09被引 1

研究自回归模型在科学计算中的泛化与记忆问题,揭示其信息传播机制缺陷。

Generalization vs. Memorization in Autoregressive Deep Learning: Or, Examining Temporal Decay of Gradient Coherence

  • 用影响函数分析模型如何吸收和传递不同物理场景的信息
  • 发现标准训练方法在长期依赖建模上存在根本性局限
  • 为改进科学模拟代理模型的设计提供可操作洞见

以自回归方式训练的基础模型作为偏微分方程(PDE)代理,有望通过超越训练范围的外推能力和少量样本下的高效微调,加速科学发现。然而,实现真正泛化——这是产生新科学洞见并确保部署鲁棒性的必要条件——仍面临重大挑战。可靠判断模型是否真正泛化而非仅记忆数据,需要能清晰区分二者的能力评估指标。本文采用影响函数形式化方法,系统刻画自回归PDE代理模型对多样化物理场景信息的吸收与传播特性,揭示了现有模型与训练流程的根本局限,并为改进代理模型设计提供了可操作的洞察。

原文摘要 · Abstract (English)

Foundation models trained as autoregressive PDE surrogates hold significant promise for accelerating scientific discovery through their capacity to both extrapolate beyond training regimes and efficiently adapt to downstream tasks despite a paucity of examples for fine-tuning. However, reliably achieving genuine generalization - a necessary capability for producing novel scientific insights and robustly performing during deployment - remains a critical challenge. Establishing whether or not these requirements are met demands evaluation metrics capable of clearly distinguishing genuine model generalization from mere memorization. We apply the influence function formalism to systematically characterize how autoregressive PDE surrogates assimilate and propagate information derived from diverse physical scenarios, revealing fundamental limitations of standard models and training routines in addition to providing actionable insights regarding the design of improved surrogates.

自回归模型PDE代理泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。