让离线视频模型变成交互式流媒体助手,只用少量自生成数据即可提升响应时机判断能力。
EvoStreaming: Your Offline Video Model Is a Natively Streaming Assistant

- 用模型自身生成数据并自动标注相关性,实现无监督的流式交互策略进化
- 仅需1000条自生成样本,流式评估得分最高提升10.8分,且不牺牲离线性能
- 适合希望低成本改造现有视频模型为实时交互系统的研究者和开发者
流媒体视频理解不仅需要观看更长视频,还要求助手在实时中决定何时回应,平衡响应速度与表达冗余。然而,多数视频-语言模型(VideoLLMs)仅针对离线推理训练,现有流媒体评测将响应时机决策外包给评估者。为此,我们提出RealStreamEval——一种逐帧多轮评估协议,使模型暴露于序列观察中,并惩罚不必要的回应。在此协议下,我们发现强健的离线VideoLLMs虽保有良好视觉理解能力,却缺乏响应时机决策策略。基于此,我们提出EvoStreaming:一种自我演化的流式适配框架,其中基础模型充当数据生成器、相关性标注者和回放策略,无需外部监督合成流式轨迹。仅使用1,000条自生成样本(比领先方法少139倍),且无需架构修改,EvoStreaming在五个开源VideoLLM骨干(Qwen2/2.5/3-VL、InternVL-3.5、MiniCPM-V4.5)上持续提升整体RealStreamEval得分最多达10.8分,同时基本保持离线视频性能。结果表明,高效的数据交互调优是将现有VideoLLMs适配为流媒体助手的可行路径。
原文摘要 · Abstract (English)
Streaming video understanding demands more than watching longer videos: assistants must decide when to speak in real time, balancing responsiveness against verbosity. Yet most video-language models (VideoLLMs) are trained for offline inference, and existing streaming benchmarks externalize this timing decision to the evaluator. We address this gap with RealStreamEval, a frame-level multi-turn evaluation protocol that exposes models to sequential observations and penalizes unnecessary responses. Under this protocol, we observed that strong offline VideoLLMs retain useful visual understanding but lack an interaction policy for deciding when to respond. Motivated by this observation, we propose EvoStreaming, a self-evolved streaming adaptation framework in which the base model itself acts as data generator, relevance annotator, and roll-out policy to synthesize streaming trajectories without external supervision. With only $1{,}000$ self-generated samples ($139\times$ less than the leading streaming instruction-tuning approach) and no architectural changes, EvoStreaming consistently improves the overall RealStreamEval score by up to $10.8$ points across five open VideoLLM backbones (Qwen2/2.5/3-VL, InternVL-3.5, MiniCPM-V4.5) while largely preserving offline video performance. These results suggest that data-efficient interaction tuning is a practical path for adapting existing VideoLLMs to streaming assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。