arXiv:2602.18006cs.CV2026-02

构建首个百万级水下多模态数据集,提升复杂环境追踪性能

MUOT_3M: A 3 Million Frame Multimodal Underwater Benchmark and the MUTrack Tracking Method

  • 基于SAM构建多模态转单模态追踪框架,融合视觉与语言信息
  • 在5个基准上实现最高8.40%的AUC提升,推理速度达24帧/秒
  • 适合海洋机器人、生态监测等实际应用,支持高效部署

水下目标追踪对海洋机器人、大规模生态监测和深海探索至关重要,但受限于高质量数据集稀缺。现有数据集规模小且仅含RGB图像,难以应对严重颜色失真、浑浊和低可见度等挑战。本文提出MUOT_3M,首个伪多模态水下追踪基准,包含3,030段视频(总计27.8小时)共300万帧,标注了32种追踪属性、677个细粒度类别,并同步提供RGB、增强RGB、估计深度和语言模态,经海洋生物学家验证。在此基础上,提出MUTrack,一种基于SAM的多模态转单模态追踪方法,包含视觉几何对齐、视觉-语言融合及四级知识蒸馏,将多模态知识迁移至单模态学生模型。在5个水下追踪基准上的评估表明,MUTrack相比最强基线,最高提升8.40% AUC和7.80%精度,同时保持24 FPS推理速度。MUOT_3M与MUTrack为可扩展、多模态训练但实际可用的水下追踪奠定了新基础。

原文摘要 · Abstract (English)

Underwater Object Tracking (UOT) is crucial for efficient marine robotics, large scale ecological monitoring, and ocean exploration; however, progress has been hindered by the scarcity of large, multimodal, and diverse datasets. Existing benchmarks remain small and RGB only, limiting robustness under severe color distortion, turbidity, and low visibility conditions. We introduce MUOT_3M, the first pseudo multimodal UOT benchmark comprising 3 million frames from 3,030 videos (27.8h) annotated with 32 tracking attributes, 677 fine grained classes, and synchronized RGB, estimated enhanced RGB, estimated depth, and language modalities validated by a marine biologist. Building upon MUOT_3M, we propose MUTrack, a SAM-based multimodal to unimodal tracker featuring visual geometric alignment, vision language fusion, and four level knowledge distillation that transfers multimodal knowledge into a unimodal student model. Extensive evaluations across five UOT benchmarks demonstrate that MUTrack achieves up to 8.40% higher AUC and 7.80% higher precision than the strongest SOTA baselines while running at 24 FPS. MUOT_3M and MUTrack establish a new foundation for scalable, multimodally trained yet practically deployable underwater tracking.

水下追踪多模态数据集视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。