arXiv:2507.16107stat.MEcs.LG2025-07被引 2

提出新方法处理缺失值,能应对数据不随机缺失且模式稀疏的问题。

Recursive Equations For Imputation Of Missing Not At Random Data With Sparse Pattern Support

  • 基于图形模型构建可构造的完整数据规律表征
  • 在MAR和MNAR下均有效,稀疏缺失模式无需额外假设
  • 适合高维数据中缺失模式支持不足的实证研究

处理缺失值的常见方法如MICE和Amelia通常假设数据为随机缺失(MAR),并通过参数或平滑假设推断缺失分布,即使某些缺失模式在数据中无支持也能继续推断。然而此类假设在实践中往往不成立,导致分析结果产生模型误设偏差。本文提出一种新的全数据分布表征方法,适用于缺失数据的图形模型。该表征具有构造性,可灵活用于MAR与非随机缺失(MNAR)机制的插补分布计算,并能处理缺乏支持的缺失模式。基于此,我们开发了多变量插补支持模式递归算法(MISPR),采用吉布斯采样,类似MICE,但在MAR与MNAR下均具一致性,且无需对无支持缺失模式引入额外假设。模拟实验表明,当数据为MAR时,MISPR表现接近MICE;当数据为MNAR时,其结果更优、偏差更小。该方法为高维数据中实际存在的非随机缺失与稀疏模式问题提供了更严谨的解决方案。

原文摘要 · Abstract (English)

A common approach for handling missing values in data analysis pipelines is multiple imputation via software packages such as MICE (Van Buuren and Groothuis-Oudshoorn, 2011) and Amelia (Honaker et al., 2011). These packages typically assume the data are missing at random (MAR), and impose parametric or smoothing assumptions upon the imputing distributions in a way that allows imputation to proceed even if not all missingness patterns have support in the data. Such assumptions are unrealistic in practice, and induce model misspecification bias on any analysis performed after such imputation. In this paper, we provide a principled alternative. Specifically, we develop a new characterization for the full data law in graphical models of missing data. This characterization is constructive, is easily adapted for the calculation of imputation distributions for both MAR and MNAR (missing not at random) mechanisms, and is able to handle lack of support for certain patterns of missingness. We use this characterization to develop a new imputation algorithm -- Multivariate Imputation via Supported Pattern Recursion (MISPR) -- which uses Gibbs sampling, by analogy with the Multivariate Imputation with Chained Equations (MICE) algorithm, but which is consistent under both MAR and MNAR settings, and is able to handle missing data patterns with no support without imposing additional assumptions beyond those already imposed by the missing data model itself. In simulations, we show MISPR obtains comparable results to MICE when data are MAR, and superior, less biased results when data are MNAR. Our characterization and imputation algorithm based on it are a step towards making principled missing data methods more practical in applied settings, where the data are likely both MNAR and sufficiently high dimensional to yield missing data patterns with no support at available sample sizes.

缺失数据插补算法非随机缺失高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。