让AI从视频手势中理解指令,实现实时人机互动。
ViSpeak: Visual Instruction Feedback in Streaming Videos
- 提出视觉指令反馈新任务,让模型从视频中识别手势等视觉指令。
- 构建了包含7个子任务的ViSpeak-Instruct数据集和评估基准。
- ViSpeak模型达到GPT-4o级别性能,适用于实时交互场景研究。
当前大型多模态模型(LMMs)主要聚焦于离线视频理解。然而,流式视频理解因时间敏感、多模态融合与交互性强,对现有模型构成巨大挑战。本文从新视角拓展流式视频理解,提出一项名为「视觉指令反馈」的新任务:模型需感知视觉内容并从中提取指令。例如,当用户挥手示意时,智能体应识别该动作并启动欢迎对话。通过视觉模态执行指令可显著提升人机交互体验。为此,我们定义了七个与视觉密切相关的子任务,并构建了用于训练的ViSpeak-Instruct数据集及用于评估的ViSpeak-Bench。进一步,提出ViSpeak模型,其在多个流式视频理解基准上达到SOTA水平,性能接近GPT-4o。经在ViSpeak-Instruct数据集微调后,该模型具备基础视觉指令反馈能力,为后续研究提供坚实基线。
原文摘要 · Abstract (English)
Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive characteristics. In this work, we aim to extend the streaming video understanding from a new perspective and propose a novel task named Visual Instruction Feedback in which models should be aware of visual contents and learn to extract instructions from them. For example, when users wave their hands to agents, agents should recognize the gesture and start conversations with welcome information. Thus, following instructions in visual modality greatly enhances user-agent interactions. To facilitate research, we define seven key subtasks highly relevant to visual modality and collect the ViSpeak-Instruct dataset for training and the ViSpeak-Bench for evaluation. Further, we propose the ViSpeak model, which is a SOTA streaming video understanding LMM with GPT-4o-level performance on various streaming video understanding benchmarks. After finetuning on our ViSpeak-Instruct dataset, ViSpeak is equipped with basic visual instruction feedback ability, serving as a solid baseline for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。