arXiv:2601.03729cs.CV2026-01被引 1

通过融合环境上下文与分类层级信息,提升海洋生物细粒度识别精度。

MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species

  • 利用目标区域引导的多尺度环境注意力建模周边场景
  • 在训练中引入分类层级辅助分类器,实现层次一致表征
  • 在多个数据集上显著优于现有方法,适合自动化海洋监测

准确的海洋生物细粒度识别对基于水下图像的可扩展生物多样性监测与生态评估至关重要。现有方法主要关注目标外观,较少利用周围环境线索和生物分类层级信息。本文提出多上下文注意力与分类层级感知网络(MATANet),用于感兴趣区域(ROI)引导的海洋生物识别。MATANet包含两个互补模块:多上下文环境注意力模块以ROI表示为查询,聚合多尺度、以ROI为中心的上下文视图中的空间块特征,实现目标条件化的环境建模;层级辅助分类器在训练中引入更高分类层级信息,促进层次一致表征,不改变最终细粒度标签预测空间或推理流程。在官方FathomNet 2025私有测试集上,使用基础骨干网络时达到1.570的层次距离,大骨干网络达1.423,显著优于最强基线值2.603。在FishCLEF2015上,准确率达0.793,层次距离1.120,优于基线0.766和1.327。消融实验表明,环境信息提供超越重复多尺度观察的补充证据,且目标条件聚合优于直接多视图拼接。在FAIR1M v2.0上的额外实验验证了该设计在非水下图像中的适用性。后检测评估中,MATANet将检测器生成的匹配ROI上细粒度分类准确率从0.828提升至0.959,支持其在自动化海洋监测中的工程应用。

原文摘要 · Abstract (English)

Accurate fine-grained recognition of marine organisms is important for scalable biodiversity monitoring and ecological assessment using underwater imagery. However, existing methods mainly focus on target appearance and make limited use of surrounding environmental cues and biological taxonomy. We propose the Multi-Context Attention and Taxonomy-Aware Network (MATANet) for region-of-interest (ROI)-guided marine organism recognition. MATANet contains two complementary components. The Multi-Context Environmental Attention Module uses the ROI representation as a query to aggregate spatial patch features from ROI-centered contextual views at multiple scales, enabling target-conditioned modeling of the surrounding environment. Level-wise auxiliary classifiers further incorporate higher taxonomic ranks during training, encouraging hierarchically consistent representations without changing the finest-label prediction space or inference procedure. On the official FathomNet 2025 Private test split, MATANet achieves a hierarchical distance of 1.570 with the base backbone and 1.423 with the large backbone, substantially outperforming the strongest benchmark value of 2.603. On FishCLEF2015, MATANet achieves an accuracy of 0.793 and a hierarchical distance of 1.120, outperforming the strongest benchmark values of 0.766 and 1.327, respectively. Ablation studies show that surrounding scene information provides complementary evidence beyond repeated multi-scale observations of the target and that target-conditioned aggregation outperforms direct multi-view concatenation. Additional experiments on FAIR1M v2.0 examine the applicability of the proposed design beyond underwater imagery. In post-detection evaluation, MATANet improves fine-grained classification accuracy on matched detector-generated ROIs from 0.828 to 0.959 supporting its engineering applicability to automated marine monitoring.

细粒度识别水下图像注意力机制分类层级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。