arXiv:2512.00252stat.MLcs.LG2025-12中稿 · ICML被引 2

用生成模型实现更精准的动态数据融合,尤其适合复杂非线性系统。

DAISI: Data Assimilation with Inverse Sampling using Stochastic Interpolants

  • 基于生成模型逆采样融合预报信息,再通过引导采样融入观测数据。
  • 在稀疏、噪声大、非线性的观测条件下,过滤精度显著优于传统方法。
  • 无需重训练生成先验,适配各类预测模型,适合科研与工程中的高维系统建模。

数据同化(DA)是科学与工程应用的核心,通过结合模型预测与稀疏、含噪观测来估计隐藏系统状态。经典高维DA方法(如集合卡尔曼滤波)依赖高斯近似,但在复杂动力学或非线性观测算子下失效。为此,我们提出DAISI,一种基于流模型的可扩展滤波算法,利用数据驱动的先验实现灵活的概率推断。核心思路是使用预训练的平稳生成先验,通过新型逆采样步骤先融入预报信息,再通过引导式条件采样同化观测数据。该方法可直接接入任意预报模型,无需在每次同化步骤中重新训练或微调生成先验。在挑战性非线性系统上的实验表明,当观测稀疏、噪声大且非线性时,DAISI能取得传统方法难以达到的准确滤波结果。DAISI代码已公开于https://github.com/Erik-Wikingsson/DAISI。

原文摘要 · Abstract (English)

Data assimilation (DA) is a cornerstone of scientific and engineering applications, combining model forecasts with sparse and noisy observations to estimate latent system states. Classical high-dimensional DA methods, such as the ensemble Kalman filter, rely on Gaussian approximations that are violated for complex dynamics or observation operators. To address this limitation, we introduce DAISI, a scalable filtering algorithm built on flow-based generative models that enables flexible probabilistic inference using data-driven priors. The core idea is to use a stationary, pre-trained generative prior that first incorporates forecast information through a novel inverse-sampling step, before assimilating observations via guidance-based conditional sampling. This allows us to leverage any forecasting model as part of the DA pipeline without having to retrain or fine-tune the generative prior at each assimilation step. Experiments on challenging nonlinear systems show that DAISI achieves accurate filtering results in regimes with sparse, noisy, and nonlinear observations where traditional methods struggle. The code for DAISI is available at https://github.com/Erik-Wikingsson/DAISI.

数据同化生成模型非线性系统概率推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。