arXiv:2604.04567stat.MLcs.LG2026-04

提出FLOWGEM方法,用生成模型修复非单调缺失数据。

Generative Modeling under Non-Monotone MAR Missingness via Approximate Wasserstein Gradient Flows

论文配图:Generative Modeling under Non-Monotone MAR Missingness via Approximate Wasserstein Gradient Flows
图 1 · 摘自论文原文
  • 基于Wasserstein梯度流迭代生成完整数据集
  • 在多种缺失模式下实现最优生成性能
  • 适合需要理论保证的高可靠数据分析场景

数据科学中缺失值普遍存在,严重威胁后续分析可靠性。尽管研究众多,针对一般非单调缺失机制的严谨非参数方法仍十分稀缺,现有方法多为经验性插补,难以确保真实分布恢复。本文提出FLOWGEM,一种从缺失随机(MAR)数据中生成完整数据集的系统性迭代方法。受忽略最大似然估计收敛性启发,该方法最小化观测数据分布与生成样本分布之间的期望KL散度,覆盖不同缺失模式。通过离散化的粒子演化实现对应的Wasserstein梯度流,速度场由局部线性密度比估计器近似。该构造形成一个迭代运输初始粒子集至目标分布的生成过程。模拟实验和真实数据基准测试表明,FLOWGEM在多种设置下表现优异,尤其在挑战性的非单调MAR机制下亦能保持领先。结果证明FLOWGEM是现有插补方法的严谨且实用替代方案,推动了理论严谨性与实证性能之间的差距缩小。

原文摘要 · Abstract (English)

The prevalence of missing values in data science poses a substantial risk to any further analyses. Despite a wealth of research, principled nonparametric methods to deal with general non-monotone missingness are still scarce. Instead, ad-hoc imputation methods are often used, for which it remains unclear whether the correct distribution can be recovered. In this paper, we propose FLOWGEM, a principled iterative method for generating a complete dataset from a dataset with values Missing at Random (MAR). Motivated by convergence results of the ignoring maximum likelihood estimator, our approach minimizes the expected Kullback-Leibler (KL) divergence between the observed data distribution and the distribution of the generated sample over different missingness patterns. To minimize the KL divergence, we employ a discretized particle evolution of the corresponding Wasserstein Gradient Flow, where the velocity field is approximated using a local linear estimator of the density ratio. This construction yields a data generation scheme that iteratively transports an initial particle ensemble toward the target distribution. Simulation studies and real-data benchmarks demonstrate that FLOWGEM achieves state-of-the-art performance across a range of settings, including the challenging case of non-monotone MAR mechanisms. Together, these results position FLOWGEM as a principled and practical alternative to existing imputation methods, and a decisive step towards closing the gap between theoretical rigor and empirical performance.

生成模型缺失数据概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。