通过多尺度激活-选择-聚合提升细粒度鸟类识别精度
Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition
- 设计多尺度激活模块,让不同阶段学习到互异的判别特征
- 引入令牌选择机制,剔除冗余信息并保留尺度特异性特征
- 多阶段结果自适应融合,显著提升模型在多个数据集上的表现
鉴于鸟类在生态系统中的关键作用,细粒度鸟类识别(FGBR)受到越来越多关注,尤其在区分相似亚类鸟类方面。尽管基于视觉变换器(ViT)的方法通常优于卷积神经网络(CNN)方法,但研究表明,普通ViT模型的感受野有限,限制了表征丰富性且易受尺度变化影响。因此,增强现有ViT模型的多尺度能力以突破这一瓶颈具有重要意义。本文提出一种新型FGBR框架——多尺度多样线索建模(MDCM),在多尺度视觉变换器(MS-ViT)的不同阶段,采用“激活-选择-聚合”范式探索不同尺度的多样化线索。首先,提出多尺度线索激活模块,确保各阶段学习到的判别线索互不相同;其次,设计多尺度令牌选择机制,剔除冗余噪声,突出各阶段的判别性、尺度特异性线索;最后,各阶段选中的令牌独立用于鸟类识别,并通过多尺度动态聚合机制自适应融合多阶段结果以做出最终决策。定性和定量实验均验证了MDCM的有效性,在多个常用FGBR基准上超越了基于CNN和ViT的现有方法。
原文摘要 · Abstract (English)
Given the critical role of birds in ecosystems, Fine-Grained Bird Recognition (FGBR) has gained increasing attention, particularly in distinguishing birds within similar subcategories. Although Vision Transformer (ViT)-based methods often outperform Convolutional Neural Network (CNN)-based methods in FGBR, recent studies reveal that the limited receptive field of plain ViT model hinders representational richness and makes them vulnerable to scale variance. Thus, enhancing the multi-scale capabilities of existing ViT-based models to overcome this bottleneck in FGBR is a worthwhile pursuit. In this paper, we propose a novel framework for FGBR, namely Multi-scale Diverse Cues Modeling (MDCM), which explores diverse cues at different scales across various stages of a multi-scale Vision Transformer (MS-ViT) in an "Activation-Selection-Aggregation" paradigm. Specifically, we first propose a multi-scale cue activation module to ensure the discriminative cues learned at different stage are mutually different. Subsequently, a multi-scale token selection mechanism is proposed to remove redundant noise and highlight discriminative, scale-specific cues at each stage. Finally, the selected tokens from each stage are independently utilized for bird recognition, and the recognition results from multiple stages are adaptively fused through a multi-scale dynamic aggregation mechanism for final model decisions. Both qualitative and quantitative results demonstrate the effectiveness of our proposed MDCM, which outperforms CNN- and ViT-based models on several widely-used FGBR benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。