轻量级水下显著实例分割模型,实时高效且精度领先。
Aqua Boundary-Saliency Attention Module for Lightweight Underwater Salient Instance Segmentation Detection Transformer

- 设计水下边界-显著性注意力模块,融合多种水下视觉线索。
- 在4个数据集上性能超越现有方法,推理延迟仅4.31-6.34毫秒。
- 适合资源勘探与水下机器人实时感知场景使用。
水下实例分割通过像素级掩码预测与实例级区分,支持海洋资源探测、生态监测与水下机器人感知。现有基于提示或辅助模态的方法依赖大型基础模型或额外模态估计,部署效率低。本文提出轻量级水下显著实例分割检测变压器(LUSIS-DETR),核心为水下边界-显著性注意力模块(AquaBSAM)。该模块通过有界残差调制,将水下边界、对比度、衰减、色度、暗通道及中心先验线索嵌入DINOv2初始化的多尺度特征中;辅助掩码监督与小目标复制粘贴仅用于训练。在四个近期水下实例分割数据集(UIIS、UIIS10K、USIS10K、USIS16K)上评估,跨类别感知与显著实例协议均达到领先性能。NVIDIA T4 GPU上使用TensorRT半精度(FP16)测试,推理延迟为4.31–6.34毫秒,支持实时推理,复现成本低。
原文摘要 · Abstract (English)
Underwater instance segmentation integrates pixel-level mask prediction and instance-level discrimination for marine resource exploration, ecological monitoring, and underwater robotic perception. Recent prompt-based and auxiliary-modality methods improve mask quality, but their reliance on large foundation models, prompt generation, or extra modality estimation complicates efficient deployment. This work introduces Lightweight Underwater Salient Instance Segmentation Detection Transformer (LUSIS-DETR), a compact detection-transformer framework built around the Aqua Boundary-Saliency Attention Module (AquaBSAM). AquaBSAM embeds underwater boundary, contrast, attenuation, chroma, dark-channel, and center-prior cues into DINOv2-initialized multi-scale features through bounded residual modulation, while auxiliary mask supervision and small-object copy-paste are training-only. Extensive evaluation on four recent underwater instance segmentation datasets, UIIS, UIIS10K, USIS10K, and USIS16K, shows competitively leading performance against previous state-of-the-art works across category-aware and salient-instance protocols. TensorRT half-precision (FP16) benchmarking on an NVIDIA T4 graphics processing unit (GPU) achieves 4.31-6.34 milliseconds (ms) latency, supporting real-time inference under an accessible reproduction setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。