构建首个水下视觉语言模型评测基准,提升机器对水下世界的理解能力。
UVLM: Benchmarking Video Language Model for Underwater World Understanding
- 融合人工与AI协作构建水下多场景数据集,涵盖419种海洋生物。
- 设计20类任务,覆盖生物与环境的观测与变化识别,支持量化评估。
- 适合作为水下智能感知研究者、海洋机器人开发者参考基准。
近年来,大型语言模型(LLMs)在人工智能领域取得显著进展,众多基于LLM的先进工作已应用于各类场景。其中,视频语言模型(VidLMs)应用尤为广泛。然而,现有研究主要聚焦陆地场景,忽视了水下观测的高难度需求。为此,我们提出UVLM——一个通过人机协同构建的水下观测基准。为保证数据质量,我们从多角度深入设计:首先,针对水下环境特点,选取具有光照变化、水体浑浊度差异及多样视角的视频;其次,确保数据多样性,涵盖不同帧率、分辨率,包含419类海洋动物以及多种静态植物和地形;再次,任务设计上分为生物与环境两大类,每类包括内容观察与变化/动作观察,共20种任务类型;最后,设计多个挑战性评估指标,实现方法间的定量比较。在两个代表性VidLM上的实验表明,在UVLM上微调能显著提升对水下世界的理解能力,同时对现有空中场景基准(如VideoMME和Perception text)也有小幅改进。数据集与提示工程将公开发布。
原文摘要 · Abstract (English)
Recently, the remarkable success of large language models (LLMs) has achieved a profound impact on the field of artificial intelligence. Numerous advanced works based on LLMs have been proposed and applied in various scenarios. Among them, video language models (VidLMs) are particularly widely used. However, existing works primarily focus on terrestrial scenarios, overlooking the highly demanding application needs of underwater observation. To overcome this gap, we introduce UVLM, an under water observation benchmark which is build through a collaborative approach combining human expertise and AI models. To ensure data quality, we have conducted in-depth considerations from multiple perspectives. First, to address the unique challenges of underwater environments, we selected videos that represent typical underwater challenges including light variations, water turbidity, and diverse viewing angles to construct the dataset. Second, to ensure data diversity, the dataset covers a wide range of frame rates, resolutions, 419 classes of marine animals, and various static plants and terrains. Next, for task diversity, we adopted a structured design where observation targets are categorized into two major classes: biological and environmental. Each category includes content observation and change/action observation, totaling 20 distinct task types. Finally, we designed several challenging evaluation metrics to enable quantitative comparison and analysis of different methods. Experiments on two representative VidLMs demonstrate that fine-tuning VidLMs on UVLM significantly improves underwater world understanding while also showing potential for slight improvements on existing in-air VidLM benchmarks, such as VideoMME and Perception text. The dataset and prompt engineering will be released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。