AquaticCLIP让机器看懂水下世界,无需人工标注
AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis
- 用200万张水下图文对无监督训练,自动生成语义对齐
- 零样本下在分割、检测等任务上超越现有方法
- 适合海洋研究者和水下视觉算法开发者使用
保护水生生物多样性对缓解气候变化至关重要,水下场景理解在辅助海洋科学家决策中发挥关键作用。本文提出AquaticCLIP,一种专为水下场景理解设计的对比语言-图像预训练模型。该模型采用新式无监督学习框架,对齐水下图像与文本,支持分割、分类、检测和目标计数等任务。基于我们构建的200万张水下图文配对数据集(来源包括YouTube、Netflix、NatGeo等异构资源),模型在无需真实标注的情况下,丰富了现有视觉-语言模型在水下领域的表现。为微调模型,我们提出提示引导的视觉编码器,通过可学习提示逐步聚合图像块特征;同时引入视觉引导机制,将视觉上下文融入语言编码器。模型通过对比预训练损失优化,实现视觉与文本模态对齐。AquaticCLIP在多个水下计算机视觉任务的零样本设置下取得显著性能提升,兼具更强鲁棒性与可解释性,树立了水下视觉-语言应用新基准。代码与数据集已公开于GitHub。
原文摘要 · Abstract (English)
The preservation of aquatic biodiversity is critical in mitigating the effects of climate change. Aquatic scene understanding plays a pivotal role in aiding marine scientists in their decision-making processes. In this paper, we introduce AquaticCLIP, a novel contrastive language-image pre-training model tailored for aquatic scene understanding. AquaticCLIP presents a new unsupervised learning framework that aligns images and texts in aquatic environments, enabling tasks such as segmentation, classification, detection, and object counting. By leveraging our large-scale underwater image-text paired dataset without the need for ground-truth annotations, our model enriches existing vision-language models in the aquatic domain. For this purpose, we construct a 2 million underwater image-text paired dataset using heterogeneous resources, including YouTube, Netflix, NatGeo, etc. To fine-tune AquaticCLIP, we propose a prompt-guided vision encoder that progressively aggregates patch features via learnable prompts, while a vision-guided mechanism enhances the language encoder by incorporating visual context. The model is optimized through a contrastive pretraining loss to align visual and textual modalities. AquaticCLIP achieves notable performance improvements in zero-shot settings across multiple underwater computer vision tasks, outperforming existing methods in both robustness and interpretability. Our model sets a new benchmark for vision-language applications in underwater environments. The code and dataset for AquaticCLIP are publicly available on GitHub at xxx.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。