构建流式视频多模态交互评测基准,推动大模型实时理解与主动响应能力评估。
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts

- 设计M4框架,实现边看边听边生成的高效流式多模态推理。
- 包含1121个视频和2290个问题,覆盖六类子任务,聚焦实时理解与主动推理。
- 首个专为流式视频场景设计的多模态交互评测集,适合模型开发者与评测研究者。
多模态语言模型(MLLMs)如GPT-4o的快速发展推动了全模态语言模型(OmniLLMs)的兴起,这类模型旨在处理并主动响应持续的多模态数据流。然而,评估其在真实流式视频场景中的交互能力仍面临巨大挑战。本文提出OmniMMI,一个面向流式视频上下文的综合性多模态交互评测基准。OmniMMI涵盖超过1,121个视频和2,290个问题,解决现有视频评测中两个被忽视的关键问题:流式视频理解与主动推理,并覆盖六类不同子任务。此外,我们提出一种新框架——多模态多路复用建模(M4),旨在实现高效推理的流式模型,支持‘边看、边听、边生成’的协同处理能力。
原文摘要 · Abstract (English)
The rapid advancement of multi-modal language models (MLLMs) like GPT-4o has propelled the development of Omni language models, designed to process and proactively respond to continuous streams of multi-modal data. Despite their potential, evaluating their real-world interactive capabilities in streaming video contexts remains a formidable challenge. In this work, we introduce OmniMMI, a comprehensive multi-modal interaction benchmark tailored for OmniLLMs in streaming video contexts. OmniMMI encompasses over 1,121 videos and 2,290 questions, addressing two critical yet underexplored challenges in existing video benchmarks: streaming video understanding and proactive reasoning, across six distinct subtasks. Moreover, we propose a novel framework, Multi-modal Multiplexing Modeling (M4), designed to enable an inference-efficient streaming model that can see, listen while generating.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。