arXiv:2506.05806cs.CV2025-06被引 10

用扩散模型实现低延迟音频驱动虚拟人视频,实时流畅对话

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

  • 采用可变长度生成与一致性训练,缩短初始输出延迟
  • 4090显卡下384x384分辨率达78帧/秒,初始延迟仅140毫秒
  • 支持状态切换与细粒度表情控制,适合交互式虚拟人应用

基于扩散模型的虚拟人生成因其卓越表现力被广泛应用,但其高计算需求限制了在实时交互应用中的部署。本文提出一种新型音频驱动肖像视频生成框架,通过可变长度生成策略降低初始视频片段生成时间,显著提升用户体验。引入针对音-图-视频的一致性模型训练策略,实现快速少步生成;结合模型量化与流水线并行进一步加速推理。为缓解扩散过程与量化带来的稳定性下降,设计专用于长时视频生成的新推理策略,确保高保真输出下的实时性。引入类别标签实现说话、倾听、空闲状态间的无缝切换,并设计细粒度面部表情控制机制以释放模型潜力。实验表明,该方法实现低延迟、流畅且真实的双向交互。在NVIDIA RTX 4090D上,384x384分辨率下最高可达78 FPS,初始生成延迟140毫秒;512x512分辨率下为45 FPS,延迟215毫秒。

原文摘要 · Abstract (English)

Diffusion-based models have gained wide adoption in the virtual human generation due to their outstanding expressiveness. However, their substantial computational requirements have constrained their deployment in real-time interactive avatar applications, where stringent speed, latency, and duration requirements are paramount. We present a novel audio-driven portrait video generation framework based on the diffusion model to address these challenges. Firstly, we propose robust variable-length video generation to reduce the minimum time required to generate the initial video clip or state transitions, which significantly enhances the user experience. Secondly, we propose a consistency model training strategy for Audio-Image-to-Video to ensure real-time performance, enabling a fast few-step generation. Model quantization and pipeline parallelism are further employed to accelerate the inference speed. To mitigate the stability loss incurred by the diffusion process and model quantization, we introduce a new inference strategy tailored for long-duration video generation. These methods ensure real-time performance and low latency while maintaining high-fidelity output. Thirdly, we incorporate class labels as a conditional input to seamlessly switch between speaking, listening, and idle states. Lastly, we design a novel mechanism for fine-grained facial expression control to exploit our model's inherent capacity. Extensive experiments demonstrate that our approach achieves low-latency, fluid, and authentic two-way communication. On an NVIDIA RTX 4090D, our model achieves a maximum of 78 FPS at a resolution of 384x384 and 45 FPS at a resolution of 512x512, with an initial video generation latency of 140 ms and 215 ms, respectively.

虚拟人扩散模型实时生成音频驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。