arXiv:2606.09949cs.LGcs.AI2026-06

主动采样让模拟聚焦难点,提升复杂方程代理模型的可靠性。

Learning Where to Simulate: Generative Active Sampling for Online PDE Surrogate Training

论文配图:Learning Where to Simulate: Generative Active Sampling for Online PDE Surrogate Training
图 1 · 摘自论文原文
  • 用扩散模型根据预测难度动态调整参数采样分布
  • 在99%分位误差上降低超40%,整体误差波动减少近一半
  • 适合需要高鲁棒性的物理模拟场景,如流体与反应扩散

数据驱动的偏微分方程(PDE)代理模型通常依赖数值求解器生成训练数据。然而,当目标是泛化到广泛配置(如初值和物理系数)时,均匀采样常忽略具有挑战性动态的轨迹,导致预测误差高且方差大。在线训练通过实时耦合数据生成与模型训练,可动态调整求解器参数。为此,本文提出在线生成式主动采样(OGAS),一种主动学习方法,通过并行训练快速扩散模型,将代理模型输出的困难信号(如损失或不确定性)映射为配置参数。该方法主动从偏向高难度的先验中采样目标信号,持续引导数据生成向挑战区域倾斜,且不增加训练时间开销。我们在二维PDE(Kuramoto-Sivashinsky、Navier-Stokes、Gray-Scott)上验证,参数规模达308维,使用多种代理架构。结果表明,所有设置下OGAS均显著改善尾部统计性能,99%分位误差下降超40%,整体误差离散度大幅降低。尽管平均误差略有上升,但有效提升了最坏情况下的模型可靠性,计算开销几乎可忽略。

原文摘要 · Abstract (English)

Data-driven PDE surrogates are trained with data produced by numerical PDE solvers. However, when the surrogate's goal is to generalize across a wide range of PDE configurations (e.g., initial conditions and physical coefficients), generating a representative training set is non-trivial. Uniform sampling of configuration parameters often under-represents trajectories exhibiting challenging dynamics, leading to high prediction errors and large error variance in the trained surrogate. Online training, where data generation and surrogate training are coupled, offers a natural advantage by allowing solver parameters to be steered on-the-fly. To efficiently exploit this capability, we introduce Online Generative Active Sampling (OGAS), an active learning method that reactively learns the relationship between configuration parameters and surrogate performance to control the sampling distribution. OGAS trains a fast diffusion model in parallel to the surrogate to act as a conditional sampler, mapping a surrogate-derived difficulty signal (e.g., loss or uncertainty) to configuration parameters. By actively drawing target signals from a prior biased toward high difficulty, OGAS continuously steers data generation toward challenging regimes without delaying the training workflow. We evaluate OGAS across 2D PDEs with distinct challenging dynamics (Kuramoto-Sivashinsky, Navier-Stokes, Gray-Scott) and up to 308 parameters, using multiple surrogate architectures. Across all settings, OGAS consistently improves tail statistics, yielding substantial reductions in errors above the 99th percentile and overall error dispersion compared to uniform sampling. While prioritizing challenging trajectories introduces a trade-off with average error, OGAS effectively ensures worst-case reliability of trained surrogates with negligible wall-time overhead.

PDE代理主动学习扩散模型在线训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。