针对缺失值依赖隐藏信息的场景,提出新型扩散模型填补方法。
Missing Pattern Recognized Diffusion Imputation Model for Missing Not At Random

- 通过模式识别器显式建模缺失规律,引导更合理的填补过程。
- 在多模态数据上,对缺失不随机情况的填补效果优于现有方法。
- 适合处理真实世界中缺失依赖隐藏值的数据问题。
缺失数据广泛存在于时间序列和图像等领域。现实中,缺失往往与未观测值本身相关,称为缺失不随机(MNAR)。本文提出缺失模式识别扩散填补模型(PRDIM),显式捕捉缺失模式并精准填补缺失值。该模型基于期望最大化(EM)算法,迭代优化观测值与缺失掩码的联合分布似然。首先使用模式识别器近似潜在缺失规律,在每次推理中提供指导,使填补结果更符合缺失信息。大量实验表明,PRDIM在多种数据模态下均能稳定实现优异的MNAR填补性能。
原文摘要 · Abstract (English)
Missing data frequently arises across diverse domains, including time-series and image domains. In the real world, missing occurrences often depend on the unobservable values themselves, which are referred to as Missing Not at Random (MNAR). In this work, we introduce the Missing Pattern Recognized Diffusion Imputation Model (PRDIM), a novel framework that explicitly captures the missing pattern and precisely imputes unobserved values. PRDIM iteratively maximizes the likelihood of the joint distribution for observed values and missing mask under an Expectation-Maximization (EM) algorithm. In this sense, we first employ a pattern recognizer, which approximates the underlying missing pattern and provides guidance during every inference toward more plausible imputations with respect to the missing information. Through extensive experiments, we demonstrate that PRDIM consistently achieves strong imputation performance under MNAR settings across multiple data modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。