arXiv:2512.17689stat.MLcs.LG2025-12中稿 · IJCAI被引 1

研究缺失值处理对可解释模型结果可信度的影响。

Imputation Uncertainty in Interpretable Machine Learning Methods

  • 比较单次与多次插补对解释方法的影响
  • 单次插补导致方差低估,覆盖率不足
  • 多次插补才能接近理论置信水平

真实数据中缺失值普遍存在,影响可解释机器学习(IML)方法的可靠性。现有研究关注偏差问题,指出不同插补方法会导致模型解释差异,但忽略了插补过程本身带来的不确定性及其对方差和置信区间的影响。本文系统比较了不同插补方法对三种IML方法——置换特征重要性、部分依赖图和Shapley值——置信区间覆盖概率的影响。结果表明,单次插补会显著低估方差,导致置信区间覆盖率偏低;在大多数情况下,只有采用多次插补才能使覆盖概率接近名义水平(如95%),从而保证解释结果的统计可靠性。

原文摘要 · Abstract (English)

In real data, missing values occur frequently, which affects the interpretation with interpretable machine learning (IML) methods. Recent work considers bias and shows that model explanations may differ between imputation methods, while ignoring additional imputation uncertainty and its influence on variance and confidence intervals. We therefore compare the effects of different imputation methods on the confidence interval coverage probabilities of the IML methods permutation feature importance, partial dependence plots and Shapley values. We show that single imputation leads to underestimation of variance and that, in most cases, only multiple imputation is close to nominal coverage.

可解释性缺失值置信区间插补方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。