评估解释模型时,人类感知比传统指标更关键。
Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings

- 用统一框架对比8种沙普利值变体在风险场景下的表现
- 3735个案例显示:量化指标与人类理解力脱节
- 适合高风险决策系统中需警惕自动化偏见的团队
沙普利值是可解释AI的核心,但其众多变体导致评估体系碎片化,缺乏实际部署共识。尽管理论差异明确,评估仍依赖量化代理指标,而这些指标与人类实用性之间的关联尚未验证。本文采用统一的近似框架,在低延迟的风险工作流中隔离八种沙普利变体的语义差异。我们在四个风险数据集及包含专业分析师的真实欺诈检测环境中开展大规模实证评估,共完成3,735个案例审查。结果揭示根本性错配:标准量化指标(如稀疏性和忠实性)与人类感知的清晰度和决策效用无关。此外,尽管无一种变体提升客观分析表现,解释却普遍增强决策信心,表明高风险场景中存在严重的自动化偏见风险。研究提示当前评估代理无法预测下游人类影响,并为运营决策系统中的变体与指标选择提供实证指导。
原文摘要 · Abstract (English)
Shapley values are a cornerstone of explainable AI, yet their proliferation into competing formulations has created a fragmented landscape with little consensus on practical deployment. While theoretical differences are well-documented, evaluation remains reliant on quantitative proxies whose alignment with human utility is unverified. In this work, we use a unified amortized framework to isolate semantic differences between eight Shapley variants under the low-latency constraints of operational risk workflows. We conduct a large-scale empirical evaluation across four risk datasets and a realistic fraud-detection environment involving professional analysts and 3,735 case reviews. Our results reveal a fundamental misalignment: standard quantitative metrics, such as sparsity and faithfulness, are decoupled from human-perceived clarity and decision utility. Furthermore, while no formulation improved objective analyst performance, explanations consistently increased decision confidence, signaling a critical risk of automation bias in high-stakes settings. These findings suggest that current evaluation proxies are insufficient for predicting downstream human impact, and we provide evidence-based guidance for selecting formulations and metrics in operational decision systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。