arXiv:2507.23642cs.CVcs.AI2025-07中稿 · GCPR 2025

EMAT提升小物体识别与分割精度,参数量更少。

Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation

  • 用新注意力机制和可学习下采样优化小物体特征
  • 在PASCAL-5i和COCO-20i上超越现有方法,参数减少4倍以上
  • 提出新评估方式,更贴近实际标注成本场景

少样本分类与分割(FS-CS)旨在利用少量标注样本同时完成多标签分类与多类别分割。尽管当前最优方法在两项任务上表现优异,但对小物体仍存在困难。为此,本文提出高效掩码注意力变压器(EMAT),通过引入新型内存高效的掩码注意力机制、可学习的下采样策略以及参数效率增强,显著提升了分类与分割准确率,尤其改善了小物体的表现。EMAT在PASCAL-5$^i$与COCO-20$^i$数据集上优于所有现有方法,且训练参数至少减少四倍。此外,针对当前评估设置忽略高成本标注的问题,本文提出两个新评估设定,更好地反映真实应用情境。

原文摘要 · Abstract (English)

Few-shot classification and segmentation (FS-CS) focuses on jointly performing multi-label classification and multi-class segmentation using few annotated examples. Although the current state of the art (SOTA) achieves high accuracy in both tasks, it struggles with small objects. To overcome this, we propose the Efficient Masked Attention Transformer (EMAT), which improves classification and segmentation accuracy, especially for small objects. EMAT introduces three modifications: a novel memory-efficient masked attention mechanism, a learnable downscaling strategy, and parameter-efficiency enhancements. EMAT outperforms all FS-CS methods on the PASCAL-5$^i$ and COCO-20$^i$ datasets, using at least four times fewer trainable parameters. Moreover, as the current FS-CS evaluation setting discards available annotations, despite their costly collection, we introduce two novel evaluation settings that consider these annotations to better reflect practical scenarios.

少样本学习分割注意力机制模型效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。