arXiv:2504.03978cs.LGcs.AI2025-04中稿 · The 3rd World Conf…被引 5

提出V-CEM模型,让可解释AI既准又易干预。

V-CEM: Bridging Performance and Intervenability in Concept-based Models

  • 用变分推断改进概念嵌入,提升干预响应性
  • 在分布内保持高精度,在分布外干预效果接近CBM
  • 适合需要可解释且强泛化能力的AI应用

基于概念的可解释AI(C-XAI)通过中间的人类可理解概念提升模型透明度并支持人工干预。概念瓶颈模型(CBMs)在分布外(OOD)场景下干预有效,但性能不如黑盒模型;概念嵌入模型(CEMs)虽在分布内(ID)表现优异,但干预能力下降。本文提出变分概念嵌入模型(V-CEM),利用变分推断增强CEMs的干预响应性。我们在多个文本和视觉数据集上评估了模型在ID性能、ID与OOD下的干预响应性,以及新提出的概念表示一致性(CRC)指标。结果表明,V-CEM在保持CEM级ID性能的同时,在OOD设置中实现了接近CBM的干预效果,显著缩小了可解释性与泛化性能之间的差距。

原文摘要 · Abstract (English)

Concept-based eXplainable AI (C-XAI) is a rapidly growing research field that enhances AI model interpretability by leveraging intermediate, human-understandable concepts. This approach not only enhances model transparency but also enables human intervention, allowing users to interact with these concepts to refine and improve the model's performance. Concept Bottleneck Models (CBMs) explicitly predict concepts before making final decisions, enabling interventions to correct misclassified concepts. While CBMs remain effective in Out-Of-Distribution (OOD) settings with intervention, they struggle to match the performance of black-box models. Concept Embedding Models (CEMs) address this by learning concept embeddings from both concept predictions and input data, enhancing In-Distribution (ID) accuracy but reducing the effectiveness of interventions, especially in OOD scenarios. In this work, we propose the Variational Concept Embedding Model (V-CEM), which leverages variational inference to improve intervention responsiveness in CEMs. We evaluated our model on various textual and visual datasets in terms of ID performance, intervention responsiveness in both ID and OOD settings, and Concept Representation Cohesiveness (CRC), a metric we propose to assess the quality of the concept embedding representations. The results demonstrate that V-CEM retains CEM-level ID performance while achieving intervention effectiveness similar to CBM in OOD settings, effectively reducing the gap between interpretability (intervention) and generalization (performance).

可解释AI概念模型干预能力泛化性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。