arXiv:2607.07767stat.MLcs.LG2026-07

用正定核密度估计实现更准确的缺失值填补,保持数据分布一致性。

Distributionally Faithful Imputation via Positive Semi-Definite Kernel Density Estimation

论文配图:Distributionally Faithful Imputation via Positive Semi-Definite Kernel Density Estimation
图 1 · 摘自论文原文
  • 基于正定核密度估计,从缺失数据中重构联合分布。
  • 在12个数据集上表现优于主流填补方法,尤其在高维下优势明显。
  • 适合需要精准分布建模的统计推断与机器学习任务。

缺失值会损害统计推断和机器学习流程,但多数填补方法依赖启发式或严格参数假设,忽略联合数据分布。本文将缺失完全随机(MCAR)下的填补问题重新建模为从遮蔽观测中进行密度估计:拟合一个其可观测边际与真实数据完全匹配的分布。利用正定(PSD)核密度,得到具有闭式边缘的凸经验风险问题,可通过牛顿内点法求解。由此产生的PSD Impute模型可从同一拟合分布中生成单次与多次填补结果,具备统计一致性,且自适应超额风险快速收敛,对高度规则的概率分布能突破维度诅咒。在1个合成数据集和11个真实世界数据集上的初步实验显示,其分布准确性已达到主流填补方法的竞争力,展现出强大的实际应用前景。

原文摘要 · Abstract (English)

Missing values undermine statistical inference and machine learning pipelines, yet most imputation methods rely on heuristics or restrictive parametric assumptions that ignore the joint data distribution. We recast imputation under missing completely at random (MCAR) as density estimation from masked observations: estimate a distribution whose observed marginals exactly match those in the data. Leveraging positive semi definite (PSD) kernel densities we obtain a convex empirical risk problem with closed form marginals, solvable by a Newton interior point method. The resulting PSD Impute model yields both single and multiple imputations from the same fitted density, enjoys statistical consistency with fast adaptive excess risk beating the curse of dimensionality for very regular probabilities. Preliminary experiments on one synthetic and eleven real world datasets already indicate competitive distributional accuracy compared with popular imputation baselines, suggesting strong practical promise.

缺失值填补密度估计正定核统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。