arXiv:2511.03730cs.HCcs.AI2025-11中稿 · manuscript of Chap…被引 2

多数XAI评估方法夸大解释效果,真实质量需更严谨验证

Not All Explanations are Created Equal: Investigating the Pitfalls of Current XAI Evaluation

  • 用棋类教学代理对比不同解释对用户的影响
  • 发现只要有解释,用户满意度就上升,无论好坏
  • 主张应重视可操作解释,而非仅追求满意度提升

可解释人工智能(XAI)旨在通过提供模型解释增强现代AI的透明度。当前评估方法多依赖用户调研或如“保真度”等客观指标,但这些方法缺乏通用性,多数研究仅比较无解释与自研解释的效果。我们指出:任何解释相较于无解释,都可能在多数指标上表现更好,从而高估解释质量。为此,我们以代理助手教学国际象棋概念为例,验证了这一陷阱——即使解释不准确,用户满意度仍会提高。我们进一步强调应聚焦可操作性解释,并分析其适用场景。本研究呼吁未来XAI研究采用更全面的评估框架,超越单纯用户满意度,真正证明解释质量。

原文摘要 · Abstract (English)

Explainable Artificial Intelligence (XAI) aims to create transparency in modern AI models by offering explanations of the models to human users. There are many ways in which researchers have attempted to evaluate the quality of these XAI models, such as user studies or proposed objective metrics like "fidelity". However, these current XAI evaluation techniques are ad hoc at best and not generalizable. Thus, most studies done within this field conduct simple user surveys to analyze the difference between no explanations and those generated by their proposed solution. We do not find this to provide adequate evidence that the explanations generated are of good quality since we believe any kind of explanation will be "better" in most metrics when compared to none at all. Thus, our study looks to highlight this pitfall: most explanations, regardless of quality or correctness, will increase user satisfaction. We also propose that emphasis should be placed on actionable explanations. We demonstrate the validity of both of our claims using an agent assistant to teach chess concepts to users. The results of this chapter will act as a call to action in the field of XAI for more comprehensive evaluation techniques for future research in order to prove explanation quality beyond user satisfaction. Additionally, we present an analysis of the scenarios in which placebic or actionable explanations would be most useful.

XAI评估方法可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。