让AI实时主动互动,像真人一样看视频并及时回应。
Proact-VL: A Proactive VideoLLM for Real-Time AI Companions

- 设计主动响应机制,结合实时视频流与对话决策
- 在三个游戏场景中实现低延迟(<1秒)且高准确率的交互
- 适合开发实时陪伴型AI,如游戏解说或导览助手
真实世界中的智能体需要具备主动性和实时性以实现类人交互,但面临三大挑战:(1)在持续视频输入下实现低延迟推理;(2)自主决定何时回应;(3)在实时约束下控制生成内容的质量与数量。本文通过解说员和引导员两个游戏场景构建智能体,并提出大型的实时游戏基准数据集(Live Gaming Benchmark),包含三种典型场景:单人解说、协同解说和用户引导。我们提出了Proact-VL框架,将多模态语言模型转化为具备主动感知与实时交互能力的智能体。大量实验表明,Proact-VL在保持强视频理解能力的同时,显著降低响应延迟并提升交互质量,验证了其在实时交互应用中的实用性。
原文摘要 · Abstract (English)
Proactive and real-time interactive experiences are essential for human-like AI companions, yet face three key challenges: (1) achieving low-latency inference under continuous streaming inputs, (2) autonomously deciding when to respond, and (3) controlling both quality and quantity of generated content to meet real-time constraints. In this work, we instantiate AI companions through two gaming scenarios, commentator and guide, selected for their suitability for automatic evaluation. We introduce the Live Gaming Benchmark, a large-scale dataset with three representative scenarios: solo commentary, co-commentary, and user guidance, and present Proact-VL, a general framework that shapes multimodal language models into proactive, real-time interactive agents capable of human-like environment perception and interaction. Extensive experiments show Proact-VL achieves superior response latency and quality while maintaining strong video understanding capabilities, demonstrating its practicality for real-time interactive applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。