构建首个融合全局与像素级理解的视频数据集,推动多模态模型全面感知视频。
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
- 引入包含精准分割掩码的视频字幕数据集,实现语言与视觉像素级对齐。
- 支持同时评估模型在整体理解和细粒度定位上的表现,覆盖千余挑战性视频。
- 适合研究多模态理解、视觉定位及通用视频分析的学者与工程师。
近年来,多模态大语言模型(MLLMs)推动了视频理解的发展,主要聚焦于视频字幕生成和问答等高层任务。与此同时,另一些工作关注密集的像素级分割任务,如类别引导或指代式目标分割。尽管两者对实现人类级视频理解均至关重要,但长期独立发展,采用不同的基准和架构。本文提出 ViCaS,一个新数据集,包含数千个复杂视频,每段视频配有详细的人工撰写字幕,并附有多个对象的时序一致、像素精确的掩码及短语定位标注。该基准可同时评估模型在整体理解与语言引导的像素级分割能力。我们还提出了经过验证的评估指标,并设计了一种有效模型架构以应对该任务。项目页面:https://ali2500.github.io/vicas-project/
原文摘要 · Abstract (English)
Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses dense, pixel-precise segmentation tasks, which typically involve category-guided or referral-based object segmentation. Although both directions are essential for developing models with human-level video comprehension, they have largely evolved separately, with distinct benchmarks and architectures. This paper aims to unify these efforts by introducing ViCaS, a new dataset containing thousands of challenging videos, each annotated with detailed, human-written captions and temporally consistent, pixel-accurate masks for multiple objects with phrase grounding. Our benchmark evaluates models on both holistic/high-level understanding and language-guided, pixel-precise segmentation. We also present carefully validated evaluation measures and propose an effective model architecture that can tackle our benchmark. Project page: https://ali2500.github.io/vicas-project/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。