arXiv:2603.18237cs.LGcs.AI2026-03

用梯度信息优化采样,提升神经模拟器的预测精度。

Gradient-Informed Temporal Sampling Improves Rollout Accuracy in PDE Surrogate Training

  • 根据模型梯度和时间覆盖度联合优化采样策略
  • 在多个PDE系统中均降低滚动预测误差
  • 适合需要高精度模拟的科学计算场景

研究人员通常在均匀采样的数值模拟数据上训练神经模拟器。但在相同预算下,系统性采样是否能提供最有效信息?如何为神经模拟器选择训练数据以最大化滚动预测准确率,是一个基础但尚未形式化的问题。现有采样方法要么坍缩到局部高信息密度区域,要么虽保持多样性却缺乏模型特异性,性能常不优于均匀采样。为此,我们提出专为神经模拟器设计的数据采样方法——梯度感知时间采样(GITS)。GITS联合优化预训练模型的局部梯度与全局时间覆盖度,从而在模型特异性和动态信息之间实现有效平衡。相比多种采样基线,GITS选取的数据在多个PDE系统、模型主干结构和采样比例下均取得更低的滚动误差。消融实验表明两项优化目标具有必要性与互补性。此外,我们分析了GITS成功与失败的典型模式及适用的PDE系统和模型架构。

原文摘要 · Abstract (English)

Researchers train neural simulators on uniformly sampled numerical simulation data. But under the same budget, does systematically sampled data provide the most effective information? A fundamental yet unformalized problem is how to sample training data for neural simulators so as to maximize rollout accuracy. Existing data sampling methods either tend to collapse into locally high-information-density regions, or preserve diversity but remain insufficiently model-specific, often leading to performance that is no better than uniform sampling. To address this, we propose a data sampling method tailored to neural simulators, Gradient-Informed Temporal Sampling (GITS). GITS jointly optimizes pilot-model local gradients and set-level temporal coverage, thereby effectively balancing model specificity and dynamical information. Compared with multiple sampling baselines, the data selected by GITS achieves lower rollout error across multiple PDE systems, model backbones and sample ratios. Furthermore, ablation studies demonstrate the necessity and complementarity of the two optimization objectives in GITS. In addition, we analyze the successful sampling patterns of GITS as well as the typical PDE systems and model backbones on which GITS fails.

神经模拟器PDE求解采样策略梯度信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。