arXiv:2412.14231cs.CV2024-12被引 3

混合可视化方法提升ViT模型可解释性,让决策过程更透明。

ViTmiX: Vision Transformer Explainability Augmented by Mixed Visualization Methods

  • 融合多种解释技术,用几何平均混合增强结果
  • 新指标验证显示解释力显著提升,物体分割更精准
  • 适合关注AI决策透明度的研究者与开发者

视觉变换器(ViT)在图像识别任务中表现卓越,依赖自注意力机制捕捉图像长程依赖。然而其复杂结构亟需强解释性方法揭示决策过程。可解释人工智能(XAI)通过可视化手段提升模型透明度与可信度。现有基于梯度或层间相关传播(LRP)的方法虽有效但存在局限。本研究提出混合多种解释技术的协同方法,显著提升ViT可解释性。我们引入几何平均融合策略,在物体分割任务中取得明显成效。为量化解释增益,提出基于鸽巢原理的新后验可解释性度量。结果表明,优化解释方法对构建可靠XAI分割系统至关重要。

原文摘要 · Abstract (English)

Recent advancements in Vision Transformers (ViT) have demonstrated exceptional results in various visual recognition tasks, owing to their ability to capture long-range dependencies in images through self-attention mechanisms. However, the complex nature of ViT models requires robust explainability methods to unveil their decision-making processes. Explainable Artificial Intelligence (XAI) plays a crucial role in improving model transparency and trustworthiness by providing insights into model predictions. Current approaches to ViT explainability, based on visualization techniques such as Layer-wise Relevance Propagation (LRP) and gradient-based methods, have shown promising but sometimes limited results. In this study, we explore a hybrid approach that mixes multiple explainability techniques to overcome these limitations and enhance the interpretability of ViT models. Our experiments reveal that this hybrid approach significantly improves the interpretability of ViT models compared to individual methods. We also introduce modifications to existing techniques, such as using geometric mean for mixing, which demonstrates notable results in object segmentation tasks. To quantify the explainability gain, we introduced a novel post-hoc explainability measure by applying the Pigeonhole principle. These findings underscore the importance of refining and optimizing explainability methods for ViT models, paving the way to reliable XAI-based segmentations.

可解释AI视觉Transformer可视化图像分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。