arXiv:2512.21150cs.CV2025-12中稿 · The IEEE/CVF Winte…被引 1

构建首个海洋生物识别与理解多模态基准,助力生态监测自动化。

ORCA: Object Recognition and Comprehension for Archiving Marine Species

  • 设计跨模态数据集,含478种海洋生物的图像与标注
  • 在14,647张图像上验证模型,发现形态重叠导致识别难题
  • 适合海洋生态、计算机视觉交叉研究者参考

海洋视觉理解对生态系统监测与保护至关重要,可实现自动化的规模化生物调查。然而,受限于训练数据稀少及缺乏将领域特异性挑战与明确计算机视觉任务对齐的系统性框架,模型应用受到限制。为此,我们提出ORCA,一个包含478种海洋生物的多模态基准,涵盖14,647张图像,共42,217个边界框标注和22,321条专家验证的实例描述。该数据集提供细粒度的视觉与文本标注,捕捉不同物种的形态特征。为推动方法进步,我们在三个任务上评估了18个先进模型:目标检测(封闭集与开放词汇)、实例描述生成和视觉定位。结果揭示了物种多样性、形态相似性及领域特殊需求带来的核心挑战,凸显海洋理解的复杂性。ORCA因此成为推动海洋领域研究的综合性基准。项目页面:http://orca.hkustvgd.com/。

原文摘要 · Abstract (English)

Marine visual understanding is essential for monitoring and protecting marine ecosystems, enabling automatic and scalable biological surveys. However, progress is hindered by limited training data and the lack of a systematic task formulation that aligns domain-specific marine challenges with well-defined computer vision tasks, thereby limiting effective model application. To address this gap, we present ORCA, a multi-modal benchmark for marine research comprising 14,647 images from 478 species, with 42,217 bounding box annotations and 22,321 expert-verified instance captions. The dataset provides fine-grained visual and textual annotations that capture morphology-oriented attributes across diverse marine species. To catalyze methodological advances, we evaluate 18 state-of-the-art models on three tasks: object detection (closed-set and open-vocabulary), instance captioning, and visual grounding. Results highlight key challenges, including species diversity, morphological overlap, and specialized domain demands, underscoring the difficulty of marine understanding. ORCA thus establishes a comprehensive benchmark to advance research in marine domain. Project Page: http://orca.hkustvgd.com/.

海洋生态多模态目标检测实例描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。