构建首个面向海洋生物的视频理解数据集,支持精准分割与语义描述。
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
- 提出两阶段海洋视频描述框架,结合视频、文本与分割掩码
- 通过视频拆分捕捉关键物体变化,提升描述语义丰富度
- 适合海洋生态研究与视频生成方向的研究者使用
海洋视频因海洋生物动态、环境复杂、摄像机运动及水下场景多样而难以理解。现有视频描述数据集多聚焦通用或人类中心场景,难以泛化至复杂的海洋环境。为此,我们提出一种面向海洋生物的两阶段视频描述流水线,构建了包含视频、文本和分割掩码三元组的综合性视频理解基准,推动海洋视频理解、分析与生成。此外,我们验证了视频拆分在检测显著物体过渡与场景变化中的有效性,显著提升了描述内容的语义信息。数据集与代码已公开于 https://msc.hkustvgd.com。
原文摘要 · Abstract (English)
Marine videos present significant challenges for video understanding due to the dynamics of marine objects and the surrounding environment, camera motion, and the complexity of underwater scenes. Existing video captioning datasets, typically focused on generic or human-centric domains, often fail to generalize to the complexities of the marine environment and gain insights about marine life. To address these limitations, we propose a two-stage marine object-oriented video captioning pipeline. We introduce a comprehensive video understanding benchmark that leverages the triplets of video, text, and segmentation masks to facilitate visual grounding and captioning, leading to improved marine video understanding and analysis, and marine video generation. Additionally, we highlight the effectiveness of video splitting in order to detect salient object transitions in scene changes, which significantly enrich the semantics of captioning content. Our dataset and code have been released at https://msc.hkustvgd.com.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。