用真实任务检验解释效果,发现用户满意不等于理解提升。
Beyond Satisfaction: From Placebic to Actionable Explanations For Enhanced Understandability
- 设计三种解释方案:无解释、形式化解释、可操作解释。
- 可操作解释组在心智模型测试中显著优于其他组。
- 适合关注解释真实作用的研究者和产品设计者。
可解释人工智能(XAI)为提升机器学习系统的透明度与可信度提供了有效工具。然而,当前对解释效果的评估多依赖主观用户调查,可能无法准确反映解释的实际效能。本文批判了过度依赖用户满意度指标的现象,探讨其能否区分有意义(可操作)与空洞(形式化)的解释。在最优社会保障退休年龄选择任务中,参与者采用三种协议:无解释、形式化解释、可操作解释。结果显示,获得可操作解释的参与者在客观心智模型测验中表现显著更优,但用户对形式化与可操作解释的满意度无明显差异。这表明仅靠主观评价无法判断解释是否真正促进用户理解。建议未来评估应结合客观任务表现与主观感受,以更准确衡量解释质量。代码见 https://github.com/Shymkis/social-security-explainer。
原文摘要 · Abstract (English)
Explainable AI (XAI) presents useful tools to facilitate transparency and trustworthiness in machine learning systems. However, current evaluations of system explainability often rely heavily on subjective user surveys, which may not adequately capture the effectiveness of explanations. This paper critiques the overreliance on user satisfaction metrics and explores whether these can differentiate between meaningful (actionable) and vacuous (placebic) explanations. In experiments involving optimal Social Security filing age selection tasks, participants used one of three protocols: no explanations, placebic explanations, and actionable explanations. Participants who received actionable explanations significantly outperformed the other groups in objective measures of their mental model, but users rated placebic and actionable explanations as equally satisfying. This suggests that subjective surveys alone fail to capture whether explanations truly support users in building useful domain understanding. We propose that future evaluations of agent explanation capabilities should integrate objective task performance metrics alongside subjective assessments to more accurately measure explanation quality. The code for this study can be found at https://github.com/Shymkis/social-security-explainer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。