arXiv:2410.11896cs.HCcs.AI2024-10被引 8

通过用户任务表现评估AI解释的有用性,更真实反映其对决策的帮助。

Study on the Helpfulness of Explainable Artificial Intelligence

  • 用用户完成代理任务的表现衡量解释是否真正有用
  • 不同XAI方法在建立信任和判断正确性上效果差异明显
  • 适合关注人机协作与可解释性落地的研究者

可解释人工智能(XAI)在医疗诊断、自动驾驶等关键领域至关重要。尽管法律、商业和伦理需求推动使用有效XAI,但方法众多且解释高度依赖上下文,仅靠自动化评估无法捕捉人类理解能力。本文提出通过用户完成代理任务的表现来评估XAI的实用性,良好表现表明解释提供了有价值信息。基于对前沿方法的用户研究,结果表明不同方法在引发信任或怀疑、正确判断AI决策方面存在显著差异。研究建议采用此方法开展以目标为导向的人本评估,实现端到端的XAI性能衡量。

原文摘要 · Abstract (English)

Explainable Artificial Intelligence (XAI) is essential for building advanced machine learning-powered applications, especially in critical domains such as medical diagnostics or autonomous driving. Legal, business, and ethical requirements motivate using effective XAI, but the increasing number of different methods makes it challenging to pick the right ones. Further, as explanations are highly context-dependent, measuring the effectiveness of XAI methods without users can only reveal a limited amount of information, excluding human factors such as the ability to understand it. We propose to evaluate XAI methods via the user's ability to successfully perform a proxy task, designed such that a good performance is an indicator for the explanation to provide helpful information. In other words, we address the helpfulness of XAI for human decision-making. Further, a user study on state-of-the-art methods was conducted, showing differences in their ability to generate trust and skepticism and the ability to judge the rightfulness of an AI decision correctly. Based on the results, we highly recommend using and extending this approach for more objective-based human-centered user studies to measure XAI performance in an end-to-end fashion.

可解释AI人机交互用户研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。