让视频大模型能主动判断何时回应,提升实时交互体验。
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
- 通过多轮强化学习,让模型自主决定是否回应
- 在5.2万段视频上训练,响应时机更准更及时
- 适合需要实时互动的场景,如智能客服、教育助手
视频多模态大语言模型在视频理解与跨模态交互方面取得了显著进展。现有系统多为回合制,需用户先发言后模型回应;而能在视频播放中主动判断回复时机,对实时应用极具潜力但挑战重重。本文提出一种全新的文本到文本主动交互方法,模型基于对话历史与当前视频帧的视觉上下文,自主决定是否回应或保持沉默。为克服以往方法需人工设定回复阈值、标注精确回复时间等问题,我们设计了一种多轮强化学习训练策略,无需精确回复时间标注即可实现及时准确的响应。我们在包含5.2万段视频的双类型对话数据集上,通过SFT与RL联合训练模型MMDuet2。实验表明,该模型在响应时机与质量上均超越现有主动式视频多模态模型,在ProactiveVideoQA基准上达到领先水平。
原文摘要 · Abstract (English)
Recent advances in video multimodal large language models (Video MLLMs) have significantly enhanced video understanding and multi-modal interaction capabilities. While most existing systems operate in a turn-based manner where the model can only reply after user turns, proactively deciding when to reply during video playback presents a promising yet challenging direction for real-time applications. In this work, we propose a novel text-to-text approach to proactive interaction, where the model autonomously determines whether to respond or remain silent at each turn based on dialogue history and visual context up to current frame of an streaming video. To overcome difficulties in previous methods such as manually tuning response decision thresholds and annotating precise reply times, we introduce a multi-turn RL based training method that encourages timely and accurate responses without requiring precise response time annotations. We train our model MMDuet2 on a dataset of 52k videos with two types of dialogues via SFT and RL. Experimental results demonstrate that MMDuet2 outperforms existing proactive Video MLLM baselines in response timing and quality, achieving state-of-the-art performance on the ProactiveVideoQA benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。