arXiv:2602.00844stat.MLcs.LG2026-02

针对多变量时间序列缺失数据,提出抗分布偏移的鲁棒填补方法。

Multivariate Time Series Data Imputation via Distributionally Robust Regularization

  • 在Wasserstein模糊集内优化最坏情况分布差异,兼顾重建误差与分布鲁棒性。
  • 在多个真实数据集上,相比基线方法显著提升填补精度,尤其在系统性缺失场景下。
  • 适合处理非平稳、缺失模式复杂的工业或医疗时间序列数据。

多变量时间序列填补常受观测数据与真实数据分布不匹配的影响,这种偏差源于时间序列非平稳性和系统性缺失的共同作用。传统方法通过点对点重建或直接分布对齐,容易过拟合有偏观测。本文提出分布鲁棒正则化填补目标(DRIO),联合最小化重建误差与在Wasserstein模糊集内,填补分布与数据分布之间的最坏情况差异。我们推导出可计算的上界代理,将无限维测度优化转化为对样本轨迹的对抗搜索,并设计了适配现代深度学习骨干网络的交替学习算法。在多种真实世界数据集上的全面实验表明,DRIO能持续提供鲁棒的填补效果,并在不同缺失情景下改善下游预测性能。

原文摘要 · Abstract (English)

Multivariate time series imputation is often compromised by mismatch between the observed and true data distributions, a bias induced by the combined effects of time-series non-stationarity and systematic missingness. Standard methods that encourage point-wise reconstruction or direct distributional alignment may overfit these biased observations. We propose the Distributionally Robust Regularized Imputer Objective (DRIO), which jointly minimizes reconstruction error and the worst-case divergence between the imputer distribution and data distributions within a Wasserstein ambiguity set. We derive a tractable upper-bound surrogate that reduces infinite-dimensional optimization over measures to adversarial search over sample trajectories, and develop an alternating learning algorithm compatible with modern deep learning backbones. Comprehensive experiments on diverse real-world datasets show that DRIO consistently provides robust imputation and suggests improved downstream forecasting under various missingness scenarios.

时间序列数据填补分布鲁棒Wasserstein

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。