arXiv:2501.09221cs.CVcs.LG2025-01IJCAI被引 1

让视觉Transformer更懂人类概念,提升可解释性与精度。

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers

  • 基于注意力机制融合多尺度特征与图像块表示
  • 在5个数据集上实现更高准确率和稳定概念解释
  • 适合需要可解释性的医疗、生物等高风险应用

随着视觉Transformer在敏感视觉任务中的广泛应用,对模型可解释性的需求日益增长。为此,研究者尝试将模型与精心标注的抽象语义概念进行前向对齐,使概念成为模型预测的全局依据,便于领域专家快速理解或干预。现有方法多为不依赖模型结构的通用插件式解释模块,未考虑基础模型(如归纳偏置、尺度不变性)在训练中的内在特性。针对此问题,本文提出ASCENT-ViT,一种基于注意力的概念学习框架,能从多尺度特征金字塔中构建尺度感知表示,并结合视觉变压器的图像块表示生成位置感知表示。这些表示通过注意力矩阵与概念标注对齐,同时捕捉空间与全局语义信息。ASCENT-ViT可作为标准ViT主干网络的分类头使用,在五个数据集上表现优异,包括三个常用基准数据集(CUB、Pascal APY、Concept-MNIST)及两个真实世界数据集(AWA2、KITS),显著提升了预测性能与概念解释的准确性与鲁棒性。

原文摘要 · Abstract (English)

As Vision Transformers (ViTs) are increasingly adopted in sensitive vision applications, there is a growing demand for improved interpretability. This has led to efforts to forward-align these models with carefully annotated abstract, human-understandable semantic entities - concepts. Concepts provide global rationales to the model predictions and can be quickly understood/intervened on by domain experts. Most current research focuses on designing model-agnostic, plug-and-play generic concept-based explainability modules that do not incorporate the inner workings of foundation models (e.g., inductive biases, scale invariance, etc.) during training. To alleviate this issue for ViTs, in this paper, we propose ASCENT-ViT, an attention-based, concept learning framework that effectively composes scale and position-aware representations from multiscale feature pyramids and ViT patch representations, respectively. Further, these representations are aligned with concept annotations through attention matrices - which incorporate spatial and global (semantic) concepts. ASCENT-ViT can be utilized as a classification head on top of standard ViT backbones for improved predictive performance and accurate and robust concept explanations as demonstrated on five datasets, including three widely used benchmarks (CUB, Pascal APY, Concept-MNIST) and 2 real-world datasets (AWA2, KITS).

视觉Transformer可解释性概念学习注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。