arXiv:2511.07923cs.CVcs.AI2025-11被引 17

无需额外训练,让模型直接识别水下生物

Exploring the Underwater World Segmentation without Extra Training

  • 用几何先验和语义对齐,将陆地视觉语言模型迁移到水下场景
  • 在255类、超2万张图像的水下数据集上实现高效分割
  • 适合做水下生态监测的科研与环保人员快速部署

精准分割海洋生物对生物多样性监测和生态评估至关重要,但现有数据集和模型仍主要聚焦陆地场景。为此,我们提出首个大规模细粒度水下分割数据集AquaOV255,包含255个类别和超过20,000张图像,覆盖多样化物种以支持开放词汇(OV)评估。同时建立首个水下开放词汇分割基准UOVSBench,整合AquaOV255与五个额外水下数据集,实现全面评估。我们还提出Earth2Ocean——一种无需额外训练的开放词汇分割框架,通过迁移陆地视觉-语言模型(VLMs)实现水下域适应。该框架包含两个核心组件:基于几何先验的视觉掩码生成器(GMG),提升局部结构感知;以及类别-视觉语义对齐模块(CSA),通过多模态大语言模型推理和场景感知模板构建增强文本嵌入。在UOVSBench上的大量实验表明,Earth2Ocean在平均性能上显著提升,且推理效率高。

原文摘要 · Abstract (English)

Accurate segmentation of marine organisms is vital for biodiversity monitoring and ecological assessment, yet existing datasets and models remain largely limited to terrestrial scenes. To bridge this gap, we introduce \textbf{AquaOV255}, the first large-scale and fine-grained underwater segmentation dataset containing 255 categories and over 20K images, covering diverse categories for open-vocabulary (OV) evaluation. Furthermore, we establish the first underwater OV segmentation benchmark, \textbf{UOVSBench}, by integrating AquaOV255 with five additional underwater datasets to enable comprehensive evaluation. Alongside, we present \textbf{Earth2Ocean}, a training-free OV segmentation framework that transfers terrestrial vision--language models (VLMs) to underwater domains without any additional underwater training. Earth2Ocean consists of two core components: a Geometric-guided Visual Mask Generator (\textbf{GMG}) that refines visual features via self-similarity geometric priors for local structure perception, and a Category-visual Semantic Alignment (\textbf{CSA}) module that enhances text embeddings through multimodal large language model reasoning and scene-aware template construction. Extensive experiments on the UOVSBench benchmark demonstrate that Earth2Ocean achieves significant performance improvement on average while maintaining efficient inference.

水下分割开放词汇零样本迁移生态监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。