arXiv:2608.05663cs.CVcs.SD2026-08

让虚拟人实时生成长视频,保持口型同步与形象一致

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

论文配图:Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
图 1 · 摘自论文原文
  • 用因果生成+自蒸馏,解决长期生成中的误差累积问题
  • 每秒生成27.12帧,超过实时播放速率,支持连续对话
  • 适合需要高质量虚拟主播的直播、教育等场景

实时长时长虚拟人音视频生成需在保证音频视觉同步和视觉一致性的同时实现因果连续合成。将预训练的双向模型适配到此场景面临两大挑战:其一,自回归地复用生成块作为上下文会引入暴露偏差,导致错误和视觉漂移随生成长度累积;其二,全局语音语句无法指示因果生成中下一步应生成的内容,尤其在仅有局部音视频上下文时。本文提出 Vorch-Streamer,一种后训练框架,可实现实时长时文本到音视频(T2AV)流式生成。构建包含80,000段12-21秒虚拟人视频的合成语料库,首先使用混合教师强制与扩散强制训练因果生成器。随后采用长时程自强制结合去噪匹配蒸馏(DMD distillation),使模型暴露于自身生成分布,同时保留预训练双向教师的质量。为显式控制语音进展,外部语言模型预测离散的25赫兹语音规划令牌,其连续特征用于条件化音频扩散分支,并对齐每个因果块与其应说内容。在有限因果上下文和四步去噪条件下,Vorch-Streamer 能以27.12帧/秒的速度联合生成音频与视频,超过24帧/秒的实时播放要求,同时保持优异的音频-口型同步与长期生成下的身份保真度。

原文摘要 · Abstract (English)

Real-time long-form avatar audio-video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting presents two key dilemmas. First, autoregressively reusing generated blocks as context creates exposure bias, causing errors and visual drift to accumulate over long rollouts. Second, a global speech utterance does not indicates a causal generator which portion should be spoken next when only limited local audio-video context is available. We present Vorch-Streamer, a post-training framework that addresses these challenges and enables real-time long-form Text-to-Audio-Video (T2AV) streaming. We construct a synthetic corpus of 80K avatar clips spanning 12-21 seconds and first train a causal generator with mixed Teacher Forcing and Diffusion Forcing. We then apply long-horizon Self Forcing with DMD distillation, exposing the model to its own rollout distribution while preserving the quality of the pretrained bidirectional teacher. To explicitly control speech progression, an external language model predicts discrete 25-Hz speech-planning tokens, whose continuous features condition the audio diffusion branch and align each causal block with the content it should speak. With bounded causal context and four-step denoising, Vorch-Streamer jointly generates audio and video from text at 27.12 FPS, exceeding the 24-FPS real-time playback rate while maintaining competitive audio-lip synchronization and strong identity preservation over long-form generation.

虚拟人生成实时生成音视频同步扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。