arXiv:2411.17991cs.CVcs.CL2024-11EMNLP被引 46

提出视频-文本二重奏交互格式,让视频大模型实时响应。

VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format

  • 用户与模型可在视频播放中任意插入文本,模拟双人对谈。
  • 在YouCook2等任务上表现显著提升,最高达90% mAP。
  • 适合直播理解、实时问答等需即时反馈的场景。

当前视频大语言模型(VideoLLM)研究多聚焦于模型架构与训练数据,而交互方式未受重视。现有方法通常以完整视频加查询为输入,模型生成回应,难以支持直播等持续播放场景,且在时间敏感任务中表现不佳。本文提出视频-文本二重奏交互格式:视频连续播放,用户与模型可随时插入文本消息,消息结束则视频继续,类似双人合奏。为此构建了MMDuetIT数据集,用于适配该交互格式,并提出多答案基准任务MAGQA以评估实时响应能力。在MMDuetIT上训练后,模型在多项时间敏感任务中表现优异:YouCook2密集视频描述任务达76% CIDEr,QVHighlights关键片段检测达90% mAP,Charades-STA时序定位达25% [email protected],且能随视频播放实时作答。

原文摘要 · Abstract (English)

Recent researches on video large language models (VideoLLM) predominantly focus on model architectures and training datasets, leaving the interaction format between the user and the model under-explored. In existing works, users often interact with VideoLLMs by using the entire video and a query as input, after which the model generates a response. This interaction format constrains the application of VideoLLMs in scenarios such as live-streaming comprehension where videos do not end and responses are required in a real-time manner, and also results in unsatisfactory performance on time-sensitive tasks that requires localizing video segments. In this paper, we focus on a video-text duet interaction format. This interaction format is characterized by the continuous playback of the video, and both the user and the model can insert their text messages at any position during the video playback. When a text message ends, the video continues to play, akin to the alternative of two performers in a duet. We construct MMDuetIT, a video-text training dataset designed to adapt VideoLLMs to video-text duet interaction format. We also introduce the Multi-Answer Grounded Video Question Answering (MAGQA) task to benchmark the real-time response ability of VideoLLMs. Trained on MMDuetIT, MMDuet demonstrates that adopting the video-text duet interaction format enables the model to achieve significant improvements in various time-sensitive tasks (76% CIDEr on YouCook2 dense video captioning, 90\% mAP on QVHighlights highlight detection and 25% [email protected] on Charades-STA temporal video grounding) with minimal training efforts, and also enable VideoLLMs to reply in a real-time manner as the video plays.

视频理解交互设计实时响应大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。