首个专为水下视觉语言理解设计的综合基准,助力海洋研究与智能探测。
UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
- 构建包含1.5万张高分辨率水下图像的多环境数据集,覆盖珊瑚礁、深海等场景。
- 含1.5万条对象指代表达和12.5万问答对,涵盖从识别到生态关系推理的任务。
- 适用于海洋科学、生态监测与水下自主探索,推动模型在复杂水下环境的泛化能力。
大型视觉语言模型(VLMs)在自然场景理解中取得显著进展,但其在水下环境中的应用仍待探索。水下图像面临光衰减严重、色彩失真、悬浮颗粒散射等挑战,且需具备海洋生态系统与生物分类学知识。为此,我们提出UWBench,一个专为水下视觉语言理解设计的综合性基准。该数据集包含15,003张高分辨率水下图像,覆盖海洋、珊瑚礁与深海等多种生境。每张图像配有经人工验证的标注:15,281条对象指代表达,精准描述海洋生物与水下结构;124,983个问答对,涵盖从物体识别到生态关系理解的多元推理能力。数据集充分反映能见度、光照条件与水体浑浊度的多样性,提供真实可信的模型评估环境。基于UWBench,我们建立了三项全面基准:详细图像描述(生成具有生态背景的场景描述)、视觉定位(精确定位海洋生物)与视觉问答(多模态推理水下环境)。对先进VLMs的广泛实验表明,水下理解仍具挑战性,改进空间巨大。本基准为推进水下视觉语言研究提供关键资源,支持海洋科学、生态监测与自主水下探索的应用。代码与基准将公开。
原文摘要 · Abstract (English)
Large vision-language models (VLMs) have achieved remarkable success in natural scene understanding, yet their application to underwater environments remains largely unexplored. Underwater imagery presents unique challenges including severe light attenuation, color distortion, and suspended particle scattering, while requiring specialized knowledge of marine ecosystems and organism taxonomy. To bridge this gap, we introduce UWBench, a comprehensive benchmark specifically designed for underwater vision-language understanding. UWBench comprises 15,003 high-resolution underwater images captured across diverse aquatic environments, encompassing oceans, coral reefs, and deep-sea habitats. Each image is enriched with human-verified annotations including 15,281 object referring expressions that precisely describe marine organisms and underwater structures, and 124,983 question-answer pairs covering diverse reasoning capabilities from object recognition to ecological relationship understanding. The dataset captures rich variations in visibility, lighting conditions, and water turbidity, providing a realistic testbed for model evaluation. Based on UWBench, we establish three comprehensive benchmarks: detailed image captioning for generating ecologically informed scene descriptions, visual grounding for precise localization of marine organisms, and visual question answering for multimodal reasoning about underwater environments. Extensive experiments on state-of-the-art VLMs demonstrate that underwater understanding remains challenging, with substantial room for improvement. Our benchmark provides essential resources for advancing vision-language research in underwater contexts and supporting applications in marine science, ecological monitoring, and autonomous underwater exploration. Our code and benchmark will be available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。