构建视频对象级理解新基准,提升模型对视频中物体的精细定位与追踪能力。
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
- 基于多智能体引擎构建70万条高质量视频指令数据
- 提出可精准捕捉时空特征的VideoRefer模型,性能优于现有方法
- 提供全面评测体系,适合研究视频细粒度理解的学者使用
视频大语言模型在通用视频理解方面表现出色,但主要聚焦整体感知,难以捕捉细粒度的空间与时间细节。同时,高质量的对象级视频指令数据稀缺,缺乏全面的评估基准也制约了其发展。为此,我们推出VideoRefer Suite,旨在增强视频大语言模型在更细粒度空间-时间层面的理解能力,即对视频中任意物体进行感知与推理。我们从数据集、模型和基准三个关键维度系统构建:首先,采用多智能体数据引擎精心构建大规模高质量的对象级视频指令数据集VideoRefer-700K;其次,提出VideoRefer模型,配备通用的时空对象编码器以提取精确的区域与序列表征;最后,精心设计VideoRefer-Bench,从多个维度全面评估视频大语言模型的空间-时间理解能力。大量实验与分析表明,我们的VideoRefer模型不仅在视频指代任务上表现优异,还提升了通用视频理解能力。
原文摘要 · Abstract (English)
Video Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and struggle with capturing fine-grained spatial and temporal details. Besides, the lack of high-quality object-level video instruction data and a comprehensive benchmark further hinders their advancements. To tackle these challenges, we introduce the VideoRefer Suite to empower Video LLM for finer-level spatial-temporal video understanding, i.e., enabling perception and reasoning on any objects throughout the video. Specially, we thoroughly develop VideoRefer Suite across three essential aspects: dataset, model, and benchmark. Firstly, we introduce a multi-agent data engine to meticulously curate a large-scale, high-quality object-level video instruction dataset, termed VideoRefer-700K. Next, we present the VideoRefer model, which equips a versatile spatial-temporal object encoder to capture precise regional and sequential representations. Finally, we meticulously create a VideoRefer-Bench to comprehensively assess the spatial-temporal understanding capability of a Video LLM, evaluating it across various aspects. Extensive experiments and analyses demonstrate that our VideoRefer model not only achieves promising performance on video referring benchmarks but also facilitates general video understanding capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。