对比多种补全方法的不确定性估计,发现准确率高不等于可信。
Beyond Accuracy: An Empirical Study of Uncertainty Estimation in Imputation
- 用多轮运行、条件采样等三种方式估算不确定性
- 发现准确率高的模型不确定性反而不准,校准误差大
- 给出选择补全工具的实用指南,适合数据清洗和下游建模
处理缺失数据是数据驱动分析的核心挑战。现代补全方法不仅追求重建精度,还关注不确定性表示与量化,但其可靠性与校准性仍不清楚。本文系统性地实证研究了补全中的不确定性,比较了三类代表性方法:统计类(MICE、SoftImpute)、分布对齐类(OT-Impute)和深度生成类(GAIN、MIWAE、TabCSDI)。实验覆盖多个数据集、缺失机制(MCAR、MAR、MNAR)和缺失率。通过多轮运行、条件采样和预测分布建模三种路径估算不确定性,并以校准曲线和期望校准误差(ECE)评估。结果表明,准确率与校准性常不一致:高重建精度模型未必提供可靠不确定性。本文分析了各类方法在准确性、校准性和运行时间间的权衡,识别出稳定配置,并为数据清洗及下游机器学习流程提供不确定性感知补全器的选择建议。
原文摘要 · Abstract (English)
Handling missing data is a central challenge in data-driven analysis. Modern imputation methods not only aim for accurate reconstruction but also differ in how they represent and quantify uncertainty. Yet, the reliability and calibration of these uncertainty estimates remain poorly understood. This paper presents a systematic empirical study of uncertainty in imputation, comparing representative methods from three major families: statistical (MICE, SoftImpute), distribution alignment (OT-Impute), and deep generative (GAIN, MIWAE, TabCSDI). Experiments span multiple datasets, missingness mechanisms (MCAR, MAR, MNAR), and missingness rates. Uncertainty is estimated through three complementary routes: multi-run variability, conditional sampling, and predictive-distribution modeling, and evaluated using calibration curves and the Expected Calibration Error (ECE). Results show that accuracy and calibration are often misaligned: models with high reconstruction accuracy do not necessarily yield reliable uncertainty. We analyze method-specific trade-offs among accuracy, calibration, and runtime, identify stable configurations, and offer guidelines for selecting uncertainty-aware imputers in data cleaning and downstream machine learning pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。