arXiv:2509.03176cs.LG2025-09

提出无阈值评估方法,解决图像解释模型评价中的误判问题。

Systematic Evaluation of Attribution Methods: Eliminating Threshold Bias and Revealing Method-Dependent Performance Patterns

  • 用AUC-IoU替代单一阈值,全面评估解释方法性能。
  • 发现单阈值评估可使方法排名波动超200个百分点。
  • 揭示不同病变尺寸下方法表现差异达269%,指导医疗影像选择。

Attribution方法通过识别关键输入特征来解释神经网络预测,但现有评估存在阈值选择偏差,可能颠倒方法排名并导致错误结论。当前协议在单一阈值下二值化归因图,阈值选择本身可使排名变化超过200个百分点。本文提出无阈值框架,采用交并比的曲线下面积(AUC-IoU)捕捉归因质量在全阈值范围内的表现。在皮肤科影像上评估七种归因方法,发现单阈值指标结果矛盾,而无阈值评估能可靠区分方法优劣。XRAI相比LIME提升31%,相比原始集成梯度提升204%;按病灶大小分层分析显示性能差异高达269%。研究确立了消除评估偏差的方法学标准,支持基于证据的归因方法选择。该框架既提供对归因行为的理论洞察,也为医疗影像等领域的稳健比较提供实用指导。

原文摘要 · Abstract (English)

Attribution methods explain neural network predictions by identifying influential input features, but their evaluation suffers from threshold selection bias that can reverse method rankings and undermine conclusions. Current protocols binarize attribution maps at single thresholds, where threshold choice alone can alter rankings by over 200 percentage points. We address this flaw with a threshold-free framework that computes Area Under the Curve for Intersection over Union (AUC-IoU), capturing attribution quality across the full threshold spectrum. Evaluating seven attribution methods on dermatological imaging, we show single-threshold metrics yield contradictory results, while threshold-free evaluation provides reliable differentiation. XRAI achieves 31% improvement over LIME and 204% over vanilla Integrated Gradients, with size-stratified analysis revealing performance variations up to 269% across lesion scales. These findings establish methodological standards that eliminate evaluation artifacts and enable evidence-based method selection. The threshold-free framework provides both theoretical insight into attribution behavior and practical guidance for robust comparison in medical imaging and beyond.

归因评估医学影像无阈值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。