arXiv:2410.15952cs.AIcs.LG2024-10被引 13

实证研究不同用户如何理解AI解释,发现现有方法有局限。

User-centric evaluation of explainability of AI with and for humans: a comprehensive empirical study

  • 用39名不同背景用户测试XAI解释的可理解性
  • 基于蘑菇数据集,发现解释效果因用户专业而异
  • 适合关注AI可解释性与用户体验的研究者

本研究聚焦人中心人工智能(HCAI),开展一项以用户为中心的实证评估,考察主流可解释人工智能(XAI)算法在帮助人类理解模型决策方面的有效性。研究采用跨学科方法,结合社会科学研究范式,评估基于梯度提升分类器(XGBClassifier)生成的解释在真实用户中的可理解性。实验包含39名来自数据科学、数据可视化及领域知识三类背景的参与者,通过访谈与问答评估其对模型解释的理解程度。模型使用来自UC Irvine机器学习仓库的食用与非食用蘑菇数据集训练。结果揭示现有XAI方法存在明显局限,强调需针对不同用户群体的信息需求设计新的解释原则与评估方法。研究方法和数据已公开,可推广至多种数据类型与用户场景,推动人机协同智能研究。

原文摘要 · Abstract (English)

This study is located in the Human-Centered Artificial Intelligence (HCAI) and focuses on the results of a user-centered assessment of commonly used eXplainable Artificial Intelligence (XAI) algorithms, specifically investigating how humans understand and interact with the explanations provided by these algorithms. To achieve this, we employed a multi-disciplinary approach that included state-of-the-art research methods from social sciences to measure the comprehensibility of explanations generated by a state-of-the-art lachine learning model, specifically the Gradient Boosting Classifier (XGBClassifier). We conducted an extensive empirical user study involving interviews with 39 participants from three different groups, each with varying expertise in data science, data visualization, and domain-specific knowledge related to the dataset used for training the machine learning model. Participants were asked a series of questions to assess their understanding of the model's explanations. To ensure replicability, we built the model using a publicly available dataset from the UC Irvine Machine Learning Repository, focusing on edible and non-edible mushrooms. Our findings reveal limitations in existing XAI methods and confirm the need for new design principles and evaluation techniques that address the specific information needs and user perspectives of different classes of AI stakeholders. We believe that the results of our research and the cross-disciplinary methodology we developed can be successfully adapted to various data types and user profiles, thus promoting dialogue and address opportunities in HCAI research. To support this, we are making the data resulting from our study publicly available.

可解释性用户研究人机交互实证评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。