arXiv:2510.15398cs.CVcs.AI2025-10被引 17

首个水下开放词汇实例分割基准,解决颜色失真与语义错位问题。

MARIS: Marine Open-Vocabulary Instance Segmentation with Geometric Enhancement and Semantic Alignment

  • 利用几何先验和语义对齐增强模型鲁棒性
  • 在MARIS数据集上超越现有基线,跨域表现更优
  • 适合水下机器人、海洋监测等场景应用

现有水下实例分割方法受限于封闭词汇预测,难以识别新海洋类别。为此,我们提出首个大规模细粒度水下开放词汇(OV)分割基准MARIS,包含有限已见类别和多样未见类别。尽管开放词汇分割在自然图像中表现良好,但其迁移到水下场景时面临严重视觉退化(如颜色衰减)和因缺乏水下类别定义导致的语义错位。为此,我们提出统一框架,包含两个互补组件:几何先验增强模块(GPEM)利用稳定的局部与结构线索,在视觉退化条件下保持目标一致性;语义对齐注入机制(SAIM)通过引入领域特定先验丰富语言嵌入,缓解语义模糊,提升未见类别的识别能力。实验表明,该框架在MARIS数据集上,无论域内还是跨域设置,均持续优于现有开放词汇基线,为未来水下感知研究奠定坚实基础。

原文摘要 · Abstract (English)

Most existing underwater instance segmentation approaches are constrained by close-vocabulary prediction, limiting their ability to recognize novel marine categories. To support evaluation, we introduce \textbf{MARIS} (\underline{Mar}ine Open-Vocabulary \underline{I}nstance \underline{S}egmentation), the first large-scale fine-grained benchmark for underwater Open-Vocabulary (OV) segmentation, featuring a limited set of seen categories and diverse unseen categories. Although OV segmentation has shown promise on natural images, our analysis reveals that transfer to underwater scenes suffers from severe visual degradation (e.g., color attenuation) and semantic misalignment caused by lack underwater class definitions. To address these issues, we propose a unified framework with two complementary components. The Geometric Prior Enhancement Module (\textbf{GPEM}) leverages stable part-level and structural cues to maintain object consistency under degraded visual conditions. The Semantic Alignment Injection Mechanism (\textbf{SAIM}) enriches language embeddings with domain-specific priors, mitigating semantic ambiguity and improving recognition of unseen categories. Experiments show that our framework consistently outperforms existing OV baselines both In-Domain and Cross-Domain setting on MARIS, establishing a strong foundation for future underwater perception research.

实例分割开放词汇水下视觉语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。