提出轻量级后处理方法MCal,矫正特征重要性解释中的缺失偏差。
Missingness Bias Calibration in Feature Attribution Explanations
- 将缺失偏差视为输出空间的表层问题,而非深层表征缺陷。
- 仅微调冻结主模型输出层的线性头,即可显著降低偏差。
- 在医疗视觉、语言和表格数据上表现优于或媲美复杂方法。
主流解释方法常因缺失偏差产生不可靠的特征重要性评分,这种系统性失真源于模型在被扰动的、分布外输入下进行探测。现有解决方案将其视为深层表征缺陷,需昂贵重训练或架构修改。本文挑战此假设,证明缺失偏差可作为模型输出空间的表层伪影处理。我们提出MCal,一种轻量级后处理方法,通过微调冻结基模型输出的简单线性头来校正该偏差。令人惊讶的是,这一简单修正在跨视觉、语言与表格等多样医疗基准上持续降低缺失偏差,性能与甚至超越此前重型方法。
原文摘要 · Abstract (English)
Popular explanation methods often produce unreliable feature importance scores due to missingness bias, a systematic distortion that arises when models are probed with ablated, out-of-distribution inputs. Existing solutions treat this as a deep representational flaw that requires expensive retraining or architectural modifications. In this work, we challenge this assumption and show that missingness bias can be effectively treated as a superficial artifact of the model's output space. We introduce MCal, a lightweight post-hoc method that corrects this bias by fine-tuning a simple linear head on the outputs of a frozen base model. Surprisingly, we find this simple correction consistently reduces missingness bias and is competitive with, or even outperforms, prior heavyweight approaches across diverse medical benchmarks spanning vision, language, and tabular domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。