arXiv:2505.05467cs.CVcs.AI2025-05NeurIPS被引 49

让离线视频大模型变主动流式助手,支持连续对话与实时响应。

StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant

  • 用记忆缓冲+衰减压缩实现长时多轮理解
  • 轻量激活模块让模型持续主动响应
  • 适合需要实时交互的视频应用开发

我们提出 StreamBridge,一个简单有效的框架,可将离线视频大模型无缝转化为具备流式能力的模型。该框架解决两大核心挑战:(1)多轮实时理解能力有限;(2)缺乏主动响应机制。StreamBridge 引入(1)结合轮次衰减压缩策略的记忆缓冲,支持长上下文多轮交互;(2)解耦且轻量的激活模型,可无痛集成至现有视频大模型中,实现持续主动响应。为支撑该框架,我们构建了 Stream-IT,一个大规模流式视频理解数据集,包含交错的视频-文本序列和多样指令格式。大量实验表明,StreamBridge 显著提升离线视频大模型在各类任务中的流式理解能力,性能超越 GPT-4o、Gemini 1.5 Pro 等专有模型,同时在标准视频理解基准上表现竞争力或更优。

原文摘要 · Abstract (English)

We present StreamBridge, a simple yet effective framework that seamlessly transforms offline Video-LLMs into streaming-capable models. It addresses two fundamental challenges in adapting existing models into online scenarios: (1) limited capability for multi-turn real-time understanding, and (2) lack of proactive response mechanisms. Specifically, StreamBridge incorporates (1) a memory buffer combined with a round-decayed compression strategy, supporting long-context multi-turn interactions, and (2) a decoupled, lightweight activation model that can be effortlessly integrated into existing Video-LLMs, enabling continuous proactive responses. To further support StreamBridge, we construct Stream-IT, a large-scale dataset tailored for streaming video understanding, featuring interleaved video-text sequences and diverse instruction formats. Extensive experiments show that StreamBridge significantly improves the streaming understanding capabilities of offline Video-LLMs across various tasks, outperforming even proprietary models such as GPT-4o and Gemini 1.5 Pro. Simultaneously, it achieves competitive or superior performance on standard video understanding benchmarks.

视频理解流式交互大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。