arXiv:2411.15605cs.CV2024-11被引 2

让视觉模型的决策理由说人话,且经得起验证。

GIFT: A Framework Towards Global Interpretable Faithful Textual Explanations of Vision Classifiers

  • 用反事实图像生成+语言模型,把局部解释转成自然语言
  • 通过干预实验量化验证每条解释的因果效力,确保真实可靠
  • 能发现模型隐藏偏见,适合研究可信AI的人看

理解深度视觉模型的决策过程对其实现安全可信部署至关重要。现有可解释性方法如显著图或概念分析常存在忠实度低、范围局限或语义模糊的问题。我们提出GIFT,一种后处理框架,旨在为视觉分类器生成全局、可解释、忠实且文本化的解释。GIFT首先生成大量忠实的局部视觉反事实样本,再利用视觉-语言模型将其转化为自然语言描述的视觉变化。这些局部解释由大语言模型聚合为简洁、可读的关于模型全局决策规则的假设。关键的是,GIFT包含验证阶段,通过图像干预定量评估每条解释的因果效应,确保最终文本解释与模型真实推理过程一致。在多种数据集上,包括合成的CLEVR基准、真实世界的CelebA人脸数据集和复杂的BDD驾驶场景,GIFT不仅揭示出有意义的分类规则,还发现了模型行为背后的意外偏见和潜在概念。整体上,GIFT弥合了局部反事实推理与全局可解释性之间的差距,为视觉模型提供了因果基础的文本解释方法。

原文摘要 · Abstract (English)

Understanding the decision processes of deep vision models is essential for their safe and trustworthy deployment in real-world settings. Existing explainability approaches, such as saliency maps or concept-based analyses, often suffer from limited faithfulness, local scope, or ambiguous semantics. We introduce GIFT, a post-hoc framework that aims to derive Global, Interpretable, Faithful, and Textual explanations for vision classifiers. GIFT begins by generating a large set of faithful, local visual counterfactuals, then employs vision-language models to translate these counterfactuals into natural-language descriptions of visual changes. These local explanations are aggregated by a large language model into concise, human-readable hypotheses about the model's global decision rules. Crucially, GIFT includes a verification stage that quantitatively assesses the causal effect of each proposed explanation by performing image-based interventions, ensuring that the final textual explanations remain faithful to the model's true reasoning process. Across diverse datasets, including the synthetic CLEVR benchmark, the real-world CelebA faces, and the complex BDD driving scenes, GIFT reveals not only meaningful classification rules but also unexpected biases and latent concepts driving model behavior. Altogether, GIFT bridges the gap between local counterfactual reasoning and global interpretability, offering a principled approach to causally grounded textual explanations for vision models.

可解释性视觉模型因果推理文本解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。