将视频理解为世界与事件流的统一模型,实现低延迟实时交互。
Video = World + Event Stream

- 把视频拆解为稳定世界与动态事件流,统一建模。
- 在25帧/秒下实现约200毫秒模型响应,总延迟550毫秒。
- 适用于实时音视频交互,适合需要低延迟的智能系统。
我们提出Wan-Streamer v0.3,将原生流式交互模型重新组织为一个统一框架:视频即世界加事件流。世界指视频展开过程中相对稳定的上下文,包括环境、场景、主体、背景声学条件、语音特征等;事件流则是世界中随时间变化的内容,如场景变动、主体行为、说话声及其他声音。该框架构建了大规模真实视频上的通用预训练任务:给定世界状态和输入,预测世界如何实时移动、变化与响应。由此获得的能力可适配多种实时下游任务。我们在实时全双工音视频交互中验证该模型,其中事件流包含代理的言语及自由行为。功能上,模型的多模态理解过程类似视觉-语言-动作闭环:将多模态用户输入映射为语言形式的语音与行为动作。Wan-Streamer v0.3保持v0.2运行点:640x368分辨率,25 FPS,160毫秒流单位,模型侧响应延迟约200毫秒,总交互延迟约550毫秒,双向网络预算350毫秒。
原文摘要 · Abstract (English)
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject behavior, speech, and other sounds. This yields a general-purpose pretraining task over large amounts of real video: given a world and incoming input, predict how the world moves, changes, and responds in real time. The resulting competence can be specialized to a broad family of real-time downstream tasks. We instantiate it on real-time full-duplex audio-visual interaction, where the event stream is the agent's speech together with free-form behavior. Functionally, the model's multimodal understanding process is vision-language-action-like: it maps multimodal user input to language-form speech and behavior actions. Wan-Streamer v0.3 preserves the v0.2 operating point: 640x368 video at 25 FPS, a 160 ms streaming unit, approximately 200 ms model-side response latency, and approximately 550 ms total interaction latency under a 350 ms bidirectional network budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。