提出可自适应更新参数的视觉解释生成方法,提升模型可解释性
Attention Lattice Adapter: Visual Explanation Generation for Visual Foundation Model
- 引入注意力晶格适配器与交替周期架构,自动选择层并优化注意力区域
- 在CUB-200-2011上实现53.2点的平均交并比提升,优于基线方法
- 适用于复杂视觉基础模型,适合关注可解释性的研究者使用
本研究针对视觉基础模型中的视觉解释生成问题。现有方法因缺乏适应性,难以应用于复杂模型。为此,我们提出一种新型解释生成方法,既能生成解释,又能部分更新模型参数以增强可解释性。该方法引入两个新机制:注意力晶格适配器(ALA)和交替周期架构(AEA)。ALA无需人工选择层,提升模型适应性与可解释性;AEA每两轮迭代更新一次ALA参数,有效缓解注意力区域过小的问题。我们在两个基准数据集CUB-200-2011和ImageNet-S上进行评估,结果表明,该方法在平均交并比(IoU)、插入得分、删除得分及插入-删除得分上均优于基线方法。尤其在CUB-200-2011数据集上,最佳模型相较基线提升了53.2点的平均交并比。
原文摘要 · Abstract (English)
In this study, we consider the problem of generating visual explanations in visual foundation models. Numerous methods have been proposed for this purpose; however, they often cannot be applied to complex models due to their lack of adaptability. To overcome these limitations, we propose a novel explanation generation method in visual foundation models that is aimed at both generating explanations and partially updating model parameters to enhance interpretability. Our approach introduces two novel mechanisms: Attention Lattice Adapter (ALA) and Alternating Epoch Architect (AEA). ALA mechanism simplifies the process by eliminating the need for manual layer selection, thus enhancing the model's adaptability and interpretability. Moreover, the AEA mechanism, which updates ALA's parameters every other epoch, effectively addresses the common issue of overly small attention regions. We evaluated our method on two benchmark datasets, CUB-200-2011 and ImageNet-S. Our results showed that our method outperformed the baseline methods in terms of mean intersection over union (IoU), insertion score, deletion score, and insertion-deletion score on both the CUB-200-2011 and ImageNet-S datasets. Notably, our best model achieved a 53.2-point improvement in mean IoU on the CUB-200-2011 dataset compared with the baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。