arXiv:2506.17936cs.HCcs.AI2025-06被引 3

研究发现,人们对概念解释中的泛化与误述难以区分,影响判断准确性。

When concept-based XAI is imprecise: Do people distinguish between generalisations and misrepresentations?

  • 通过铁路安全场景实验,测试人们如何评估基于概念的AI解释
  • 泛化无关特征的解释反而比精准匹配的解释评分更低
  • 对关键特征的不准确解释极为敏感,但无法识别真正泛化

基于概念的可解释人工智能(C-XAI)能让人们观察到AI模型学到的表征。这在涉及高阶语义信息(如行为和关系)判断抽象类别(如危险)的任务中尤为重要。此时,模型需超越具体情境细节进行泛化,而这种能力可通过随机化无关特征的C-XAI输出体现。然而,人们是否能理解此类泛化,并将其与不良的不精确性区分开尚不清楚。本研究在模拟的铁路安全评估场景中,让参与者评估一个将含人交通场景分类为危险或非危险的AI系统表现。分类结果通过类似图像片段的概念进行解释,这些片段在匹配被分类图像时,或涉及高度相关特征(如人与轨道的关系),或涉及较不相关特征(如人的动作)。出乎意料的是,泛化无关特征的解释评分低于精确匹配的解释;其评分甚至不优于对不相关特征的系统性误述。相反,参与者对相关特征的不精确性高度敏感。该结果质疑了人们能从C-XAI输出中轻易推断模型是否具备深层理解的假设。

原文摘要 · Abstract (English)

Concept-based explainable artificial intelligence (C-XAI) can let people see which representations an AI model has learned. This is particularly important when high-level semantic information (e.g., actions and relations) is used to make decisions about abstract categories (e.g., danger). In such tasks, AI models need to generalise beyond situation-specific details, and this ability can be reflected in C-XAI outputs that randomise over irrelevant features. However, it is unclear whether people appreciate such generalisation and can distinguish it from other, less desirable forms of imprecision in C-XAI outputs. Therefore, the present study investigated how the generality and relevance of C-XAI outputs affect people's evaluation of AI. In an experimental railway safety evaluation scenario, participants rated the performance of a simulated AI that classified traffic scenes involving people as dangerous or not. These classification decisions were explained via concepts in the form of similar image snippets. The latter differed in their match with the classified image, either regarding a highly relevant feature (i.e., people's relation to tracks) or a less relevant feature (i.e., people's action). Contrary to the hypotheses, concepts that generalised over less relevant features were rated lower than concepts that matched the classified image precisely. Moreover, their ratings were no better than those for systematic misrepresentations of the less relevant feature. Conversely, participants were highly sensitive to imprecisions in relevant features. These findings cast doubts on the assumption that people can easily infer from C-XAI outputs whether AI models have gained a deeper understanding of complex situations.

可解释AI人类判断概念解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。