基于流形假设的缺失数据填补方法,兼顾几何结构与不确定性量化。
Missing Data Imputation under Manifold Hypothesis

- 利用变分自编码器挖掘数据内在低维流形结构
- 通过采样-重要性-重采样实现条件采样,填补缺失值并量化不确定性
- 支持实时填补,适合需要动态处理的数据场景
流形假设认为高维数据集中于低维嵌入流形附近。近年来,混合变分自编码器(VAEs)为忠实提取此类底层结构提供了强大工具。由此产生的几何结构自然引入变量间的局部与全局关系,从而提供系统化的缺失数据填补方式。我们提出一种基于模型的填补方法,通过采样-重要性-重采样(SIR)过程从 $ p(m{x}_{ ext{mis}} igm| m{x}_{ ext{obs}}) $ 中采样,该过程可进一步在潜在空间中结合联合扩散模型增强。所提方法在尊重数据底层几何结构的前提下,实现了与当前最优方法相当的性能,能对填补结果进行不确定性量化,且为基于模型的方法,支持无需重新运行整个流程的实时填补。
原文摘要 · Abstract (English)
The manifold hypothesis posits that high-dimensional data are concentrated near a low-dimensional embedded manifold. Recent advances in mixture variational autoencoders (VAEs) provide a powerful tool for extracting such underlying structure in a faithful manner. The resulting geometric structure naturally introduces local and global relationships among variables, thereby providing a systematic way of imputing missing data. We propose a model-based imputation method that enables sampling from \( p(\bm{x}_{\mathrm{mis}} \mid \bm{x}_{\mathrm{obs}}) \) via a sampling-importance-resampling (SIR) procedure, which can be further augmented with a joint diffusion model in the latent space. Our method imputes missing data while respecting the underlying geometry, achieves competitive performance compared to state-of-the-art procedures, quantifies uncertainty in the imputations, and is model-based, thereby enabling on-the-fly imputation without rerunning the entire procedure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。