arXiv:2510.27481cs.CV2025-10NeurIPS被引 16

构建首个大规模水下多任务数据集,提升水下场景理解模型性能

NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding

  • 基于物理成像模型设计可插拔特征增强模块,恢复水下图像清晰度
  • 在145万图文对数据上训练,显著提升基线模型在8项任务上的表现
  • 适合从事水下机器人、海洋探测的科研与工程人员参考

水下探索对资源勘探、国家安全等具有重要意义。本文研究水下场景理解方法,旨在实现自动化水下探测。该任务需多粒度、多任务感知,但缺乏大规模多任务指令微调数据集限制了进展。为此,我们构建了包含1.45百万图像-文本对的NautData数据集,支持8项水下场景理解任务,支持模型开发与全面评估。水下图像退化是普遍挑战,影响任务效果。为此,我们引入源自水下成像模型的物理先验,提出可插拔视觉特征增强(VFE)模块,显式恢复清晰水下信息。将该模块集成至LLaVA-1.5和Qwen2.5-VL等主流基线模型中,构建水下多模态模型NAUTILUS。在NautData及公开水下数据集上的实验表明,VFE模块持续提升基线模型在多数任务上的性能,验证了NAUTILUS在水下场景理解领域的优势。数据与模型开源:https://github.com/H-EmbodVis/NAUTILUS。

原文摘要 · Abstract (English)

Underwater exploration offers critical insights into our planet and attracts increasing attention for its broader applications in resource exploration, national security, etc. We study the underwater scene understanding methods, which aim to achieve automated underwater exploration. The underwater scene understanding task demands multi-task perceptions from multiple granularities. However, the absence of large-scale underwater multi-task instruction-tuning datasets hinders the progress of this research. To bridge this gap, we construct NautData, a dataset containing 1.45 M image-text pairs supporting eight underwater scene understanding tasks. It enables the development and thorough evaluation of the underwater scene understanding models. Underwater image degradation is a widely recognized challenge that interferes with underwater tasks. To improve the robustness of underwater scene understanding, we introduce physical priors derived from underwater imaging models and propose a plug-and-play vision feature enhancement (VFE) module, which explicitly restores clear underwater information. We integrate this module into renowned baselines LLaVA-1.5 and Qwen2.5-VL and build our underwater LMM, NAUTILUS. Experiments conducted on the NautData and public underwater datasets demonstrate the effectiveness of the VFE module, consistently improving the performance of both baselines on the majority of supported tasks, thus ensuring the superiority of NAUTILUS in the underwater scene understanding area. Data and models are available at https://github.com/H-EmbodVis/NAUTILUS.

水下理解多模态视觉增强大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。