用扰动法评估概念解释的可信度,让AI决策更透明可靠。
ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

- 通过扰动输入区域,检测概念响应变化来评估解释可靠性。
- MedSAM路径空间定位更强,代理模型拟合度达R²=0.8503。
- 适用于医疗影像等需高可信解释的场景,适合关注AI可解释性研究者。
基于概念的可解释人工智能能提升模型推理的人类可理解性,但概念级输出未必可信。我们提出ConceptSMILE,一种模型无关的扰动审计框架,用于评估概念解释的可靠性。该框架将SMILE的扰动逻辑从特征或区域层级扩展至人类可理解的概念解释层面,通过扰动输入区域、测量概念响应变化、引入局部权重并拟合XGBoost代理模型,以评估归因准确率、代理模型保真度、忠实性、稳定性与一致性。我们在视网膜眼底图像上对比了MedSAM生成的视觉概念与基于视觉语言模型(VLM)的语义概念。结果显示,不同概念与路径的可靠性各异:MedSAM在空间归因上表现更优,代理模型拟合度最高(R²=0.8503,R_w²=0.8465);而VLM路径在血管忠实性及特定伪影条件下的稳定性更强。ConceptSMILE为概念解释提供了独立的可信度审计层。
原文摘要 · Abstract (English)
Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework for evaluating the reliability of concept-based explanations. Rather than replacing SMILE, ConceptSMILE extends its perturbation-based logic from feature- or region-level attribution to the auditing of human-understandable concept explanations. The framework perturbs input regions, measures concept-response shifts, applies locality weighting, and fits an XGBoost surrogate to approximate local concept behaviour. Reliability is assessed through attribution accuracy, surrogate fidelity, faithfulness, stability, and consistency. We evaluate ConceptSMILE on retinal fundus images by comparing MedSAM-derived visual concepts with VLM-based semantic concepts. Results show that reliability varies across concepts and pathways: MedSAM achieves stronger spatial attribution and the highest surrogate fidelity ($R^2 = 0.8503$, $R_w^2 = 0.8465$), while the VLM pathway shows stronger vessel faithfulness and stronger stability under selected artefact conditions. ConceptSMILE provides an independent audit layer for evaluating the trustworthiness of concept-based XAI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。