arXiv:2606.01620cs.CV2026-06被引 1

用参考图和语音实时生成高清会动的肖像视频,速度远超传统方法。

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

论文配图:Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
图 1 · 摘自论文原文
  • 用因果视频变分自编码器压缩视频,结合参考图引导动态信息
  • 生成速度显著快于基线模型,质量媲美甚至超越大模型
  • 适合需要低延迟交互的数字人、虚拟主播等实时应用

视频扩散模型极大提升了肖像视频生成效果,但计算开销大,难以用于交互场景。本文提出一种基于语音和参考图像的可流式传输说话肖像视频生成框架。为适应流式传输,设计了因果视频变分自编码器(VAE)实现深度隐空间压缩,并引入自回归隐空间去噪模型。该因果VAE可灵活接入多个参考图像作为引导,使网络聚焦于动态变化而非静态外观,提升压缩效率与重建质量;同时扩展残差自编码范式以更好处理时空因果性。生成器采用修正流变换器架构,以分块自回归方式生成视频隐变量。实验表明,该方法可实现实时高质量说话肖像视频生成,速度显著优于基线模型,在真实感、生动性和视频质量上达到或超过大型模型水平。

原文摘要 · Abstract (English)

Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applications. This work presents a framework for streamable talking portrait video generation conditioned on speech audio and reference images. Designed meticulously for streaming scenarios, it features a causal video VAE for deep latent compression and an autoregressive latent denoising model. Our causal VAE integrates a variable number of reference images as guidance, allowing the network to focus on dynamic information rather than static appearance, thereby enhancing compression efficacy and reconstruction quality. Additionally, we extend the residual auto-encoding paradigm to improve spatial-temporal causality handling in our VAE. The generator is based on a Rectified Flow Transformer architecture and produces video latents in a blockwise auto-regressive manner. Our method enables the real-time generation of high-quality talking portrait videos, achieving speeds significantly faster than baseline models. Furthermore, comprehensive experiments demonstrate that it is on par with or even outperforms these large models in realism, vividness, and video quality.

视频生成扩散模型实时生成隐空间压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。