根据用户指令动态选帧,提升视频理解精度
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
- 基于指令生成关键帧选择策略,自动适应不同任务需求
- 在40K视频数据集上实现50万条精准时序定位标注
- 适配多模态视频理解任务,尤其擅长复杂指令跟随
尽管视频大语言模型在多模态理解和推理任务中展现出巨大潜力,如何高效选取视频中最信息丰富的帧仍是一个关键挑战。现有方法通过减少帧间冗余或使用无监督事件定位来优化帧采样,但在处理复杂的指令遵循任务和需要精确时序建模的场景时表现有限,导致语义对齐和时序推理性能不足。为此,我们提出视频指令时序定位框架VideoITG,旨在根据用户指令自适应定制帧采样策略。具体地,我们设计了VidThinker流水线,通过生成指令相关的描述、检索相关视频片段并选择关键帧,实现高效的监督信号构建。利用VidThinker,我们构建了包含4万视频和50万时序定位标注的VideoITG-40K数据集。我们的即插即用式VideoITG模型利用视频大语言模型的视觉-语言对齐与推理能力,实现判别性帧选择,在多个多模态视频理解基准上持续提升性能,验证了其有效性与应用潜力。
原文摘要 · Abstract (English)
While Video Large Language Models (Video-LLMs) have shown significant potential in multimodal understanding and reasoning tasks, how to efficiently select the most informative frames from videos remains a critical challenge. Existing methods attempt to optimize frame sampling by reducing inter-frame redundancy or employing unsupervised event localization. However, these approaches often fall short in handling complex instruction-following tasks and scenarios that demand precise temporal modeling, resulting in limited performance in both semantic alignment and temporal reasoning. To address the above challenges, we introduce Instructed Temporal Grounding for Videos (VideoITG), a framework aiming to adaptively customize frame sampling strategies based on user instructions. Specifically, we design the VidThinker pipeline, which automates annotation by generating instruction-conditioned captions, retrieving relevant video segments, and selecting key frames to enable efficient supervision. Using VidThinker, we build the VideoITG-40K dataset with 40K videos and 500K temporal grounding annotations. Our plug-and-play VideoITG model leverages Video-LLMs' visual-language alignment and reasoning for discriminative frame selection. VideoITG consistently boosts the performance on multiple multimodal video understanding benchmarks, demonstrating its effectiveness and potential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。