构建实时视频交互评测基准,推动视频大模型在线对话能力发展
RIVER: A Real-Time Interaction Benchmark for Video LLMs

- 设计包含回溯记忆、实时感知与主动预测的三阶段交互框架
- 实测表明离线模型在实时场景中表现显著下降,尤其缺乏长期记忆
- 提出通用改进方法,提升模型在动态对话中的灵活响应能力
多模态大语言模型虽已展现强大能力,但几乎均基于离线模式,难以支持实时交互。为此,我们提出实时视频交互评测基准 RIVER Bench,用于评估在线视频理解性能。RIVER Bench 设计了包含回溯记忆、实时感知与主动预测的任务框架,更贴近真实对话场景而非一次性回答整段视频。我们基于多样来源和不同长度的视频进行精细标注,并明确定义了实时交互格式。跨多种模型类别的评估显示,尽管离线模型在单轮问答中表现良好,但在实时处理中表现不佳,尤其在长期记忆与未来感知方面存在明显缺陷。针对现有模型在在线视频交互中的局限性,我们提出一种通用改进方法,使模型能更灵活地实现实时交互。数据集与代码已公开于 https://github.com/OpenGVLab/RIVER。
原文摘要 · Abstract (English)
The rapid advancement of multimodal large language models has demonstrated impressive capabilities, yet nearly all operate in an offline paradigm, hindering real-time interactivity. Addressing this gap, we introduce the Real-tIme Video intERaction Bench (RIVER Bench), designed for evaluating online video comprehension. RIVER Bench introduces a novel framework comprising Retrospective Memory, Live-Perception, and Proactive Anticipation tasks, closely mimicking interactive dialogues rather than responding to entire videos at once. We conducted detailed annotations using videos from diverse sources and varying lengths, and precisely defined the real-time interactive format. Evaluations across various model categories reveal that while offline models perform well in single question-answering tasks, they struggle with real-time processing. Addressing the limitations of existing models in online video interaction, especially their deficiencies in long-term memory and future perception, we proposed a general improvement method that enables models to interact with users more flexibly in real time. We believe this work will significantly advance the development of real-time interactive video understanding models and inspire future research in this emerging field. Datasets and code are publicly available at https://github.com/OpenGVLab/RIVER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。