arXiv:2603.24006cs.CV2026-03

首个大规模水下视频目标分割数据集,解决海洋探测中图像模糊与伪装难题。

UW-VOS: A Large-Scale Dataset for Underwater Video Object Segmentation

  • 构建半自动数据引擎生成1431段视频,含30万+标注掩膜。
  • 新模型SAM-U仅2%参数量,性能超越现有方法13点指标。
  • 揭示小目标、伪装、进出场景是主要挑战,适合水下视觉研究者。

水下视频目标分割(VOS)对海洋探索至关重要,但开放域方法因颜色失真、对比度低和普遍伪装而性能显著下降。主要瓶颈在于高质量训练数据的缺乏。为此,我们提出首个大规模水下VOS基准UW-VOS,包含1,431个视频序列、409个类别及309,295个掩膜标注,通过半自动数据引擎并经严格人工验证构建。我们进一步提出SAM-U,一种参数高效的框架,通过在图像编码器中插入轻量适配器,将SAM2迁移至水下领域,仅需约2%可训练参数即实现最优性能。大量实验表明,现有方法在UW-VOS上平均$/mathcal{J} ext{ extbackslash}& ext{ extbackslash} ext{F}$下降13点,而SAM-U有效弥合了域间差距。基于属性的详细分析进一步识别出小目标、伪装及进出场景为关键瓶颈,为未来鲁棒水下感知研究提供方向。

原文摘要 · Abstract (English)

Underwater Video Object Segmentation (VOS) is essential for marine exploration, yet open-air methods suffer significant degradation due to color distortion, low contrast, and prevalent camouflage. A primary hurdle is the lack of high-quality training data. To bridge this gap, we introduce $\textbf{UW-VOS}$, the first large-scale underwater VOS benchmark comprising 1,431 video sequences across 409 categories with 309,295 mask annotations, constructed via a semi-automatic data engine with rigorous human verification. We further propose $\textbf{SAM-U}$, a parameter-efficient framework that adapts SAM2 to the underwater domain. By inserting lightweight adapters into the image encoder, SAM-U achieves state-of-the-art performance with only $\sim$2$\%$ trainable parameters. Extensive experiments reveal that existing methods experience an average 13-point $\mathcal{J}\&\mathcal{F}$ drop on UW-VOS, while SAM-U effectively bridges this domain gap. Detailed attribute-based analysis further identifies small targets, camouflage, and exit-re-entry as critical bottlenecks, providing a roadmap for future research in robust underwater perception.

视频分割水下视觉数据集迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。