提升视频大模型对具体时间片段的精准定位能力
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
- 增加时间流与离散时间标记,强化时序建模
- 在多个细粒度任务上超越现有模型,最高提升12.6%
- 适合需要精确时间理解的视频分析场景
视频大语言模型在粗粒度视频理解方面表现优异,但在细粒度时间定位上仍存在困难。本文提出Grounded-VideoLLM,一种能精细感知和推理视频特定时刻的新模型。我们发现现有模型缺乏有效的时序建模与时间戳表示,为此引入两个改进:(1) 增加时间流以编码帧间关系;(2) 使用富含时间知识的离散时间标记来表示时间戳。为优化训练,采用多阶段策略,从简单的视频描述任务逐步过渡到复杂的时间定位任务。此外,通过自动化标注流程构建了一个带标注的视频问答数据集。大量实验表明,Grounded-VideoLLM在时间句定位、密集视频描述和带标注视频问答等细粒度任务中表现卓越,且具备作为通用视频助手的巨大潜力。
原文摘要 · Abstract (English)
Video Large Language Models (Video-LLMs) have demonstrated remarkable capabilities in coarse-grained video understanding, however, they struggle with fine-grained temporal grounding. In this paper, we introduce Grounded-VideoLLM, a novel Video-LLM adept at perceiving and reasoning over specific video moments in a fine-grained manner. We identify that current Video-LLMs have limitations for fine-grained video understanding since they lack effective temporal modeling and timestamp representation. In light of this, we sharpen our model by incorporating (1) an additional temporal stream to encode the relationships between frames and (2) discrete temporal tokens enriched with specific time knowledge to represent timestamps. To optimize the training of Grounded-VideoLLM, we employ a multi-stage training scheme, beginning with simple video-captioning tasks and progressively introducing video temporal grounding tasks of increasing complexity. To further enhance Grounded-VideoLLM's temporal reasoning capability, we also curate a grounded VideoQA dataset by an automatic annotation pipeline. Extensive experiments demonstrate that Grounded-VideoLLM not only excels in fine-grained grounding tasks such as temporal sentence grounding, dense video captioning, and grounded VideoQA, but also shows great potential as a versatile video assistant for general video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。