用双路注意力增强植物图像识别,有效避开背景干扰。
AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images
- 双分支结构融合原始图与分割图,通过跨视角注意力融合特征。
- 在叶基形状等任务上准确率提升,植物区域注意力对齐提高18.03个百分点。
- 适合需要精准定位植物器官的植物分类研究者使用。
从腊叶标本图像中自动识别植物性状对植物科学至关重要,但背景元素(如文字标签、固定痕迹、色卡)常导致模型依赖虚假线索而非植物形态,降低泛化性和可解释性。本文提出AT-ViT,一种双分支视觉变换器,通过多尺度、多视角交叉注意力融合机制,联合编码原始腊叶扫描图及其分割生成的对应图像。模型还引入掩码引导的块加权机制,强化植物相关区域特征,抑制背景驱动信息。在多个性状分类任务(如叶基形状、刺状物)中,AT-ViT显著提升准确率,改善注意力在植物区域的定位能力,在合成背景扰动下表现更鲁棒。相较于CrossViT,其植物区域对齐平均交并比(Avg IoU_p)提升15.66至18.03个百分点,背景重叠度(Avg IoU_b)下降27.92至31.02个百分点;在背景噪声条件下,相比ResNet101最高提升32.32分,相比CrossViT最高提升5.07分。
原文摘要 · Abstract (English)
Automated plant traits recognition from herbarium images is essential for plant sciences, yet remains challenging because background elements (e.g., textual labels, mounting artifacts, and color charts) can introduce shortcut learning, leading models to rely on spurious non-plant cues rather than plant morphology. This bias degrades both generalization and interpretability. In this paper, we introduce AT-ViT, a dual-branch Vision Transformer that jointly encodes raw herbarium scans and their segmented-derived counterparts via a multi-scale, multi-view cross-attention fusion scheme. AT-ViT further incorporates a mask-guided patch weighting mechanism that amplifies plant-relevant regions and attenuates background-driven features. By learning from the original scans while being guided by segmentation masks through the mask-guided patch reweighting mechanism, the model is encouraged to focus on plant organs and learn plant-centric representations more effectively. Across multiple trait classification tasks (e.g., leaf base shape, thorns), AT-ViT delivers consistent accuracy gains, improves attention localization on plant regions, and exhibits increased robustness under synthetic background perturbations. Specifically, AT-ViT substantially improves spatial attention grounding, boosting plant-region alignment (Avg IoU_p: +15.66 to +18.03 pp) while reducing background overlap (Avg IoU_b: -27.92 to -31.02 pp) relative to CrossViT, and remains markedly more robust to background perturbations, outperforming ResNet101 by up to +32.32 accuracy points and CrossViT by up to +5.07 points under background-noise conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。