arXiv:2507.14787cs.CVcs.AI2025-07被引 8

让视觉Transformer看清光谱数据,一键生成可解释的3D注意力图。

FOCUS: Fused Observation of Channels for Unveiling Spectra

  • 用特定波段提示和可学习[空洞]令牌,引导注意力聚焦有效光谱区。
  • 提升波段级重合度15%,减少注意力坍塌超40%,结果贴合专家标注。
  • 零梯度、不改模型,参数增加不足1%,适合实际光谱应用。

高光谱成像(HSI)捕捉数百个窄而连续的波段,广泛应用于生物、农业与环境监测。然而,现有视觉变压器(ViT)在该场景下的可解释性研究仍处于空白:一方面,现有显著性方法难以捕捉有意义的光谱线索,常将注意力集中于类别标记;另一方面,全谱段ViT因高维数据计算成本过高,难以用于可解释性分析。本文提出FOCUS,首个实现冻结ViT下高效可靠空间-光谱可解释性的框架。其核心包含两类设计:针对类别的光谱提示,引导注意力关注语义相关的波段组;以及通过吸引损失训练的可学习[空洞]令牌,吸收噪声或冗余注意力。二者协同使单次前向传播即可生成稳定且可解释的3D显著性图与光谱重要性曲线,无需反向传播或修改主干网络。实验显示,该方法提升波段级交并比15%,注意力坍塌减少超过40%,显著性结果与专家标注高度一致。仅增加不足1%参数量,使高分辨率ViT的可解释性在真实高光谱应用中成为可能,弥合了黑箱建模与可信决策之间的长期鸿沟。

原文摘要 · Abstract (English)

Hyperspectral imaging (HSI) captures hundreds of narrow, contiguous wavelength bands, making it a powerful tool in biology, agriculture, and environmental monitoring. However, interpreting Vision Transformers (ViTs) in this setting remains largely unexplored due to two key challenges: (1) existing saliency methods struggle to capture meaningful spectral cues, often collapsing attention onto the class token, and (2) full-spectrum ViTs are computationally prohibitive for interpretability, given the high-dimensional nature of HSI data. We present FOCUS, the first framework that enables reliable and efficient spatial-spectral interpretability for frozen ViTs. FOCUS introduces two core components: class-specific spectral prompts that guide attention toward semantically meaningful wavelength groups, and a learnable [SINK] token trained with an attraction loss to absorb noisy or redundant attention. Together, these designs make it possible to generate stable and interpretable 3D saliency maps and spectral importance curves in a single forward pass, without any gradient backpropagation or backbone modification. FOCUS improves band-level IoU by 15 percent, reduces attention collapse by over 40 percent, and produces saliency results that align closely with expert annotations. With less than 1 percent parameter overhead, our method makes high-resolution ViT interpretability practical for real-world hyperspectral applications, bridging a long-standing gap between black-box modeling and trustworthy HSI decision-making.

高光谱视觉变压器可解释性光谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。