让AI Avatar实时互动生成视频,速度提升20倍且画面更稳。
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
- 改进了在线策略蒸馏方法,解决多模态输入下的画面闪烁和失真问题。
- 在多个数据集上实现与原模型相当的画质,推理成本降低20倍。
- 构建了实时交互系统LiveTalk,支持多轮对话且响应延迟降至即时。
通过扩散模型实现实时视频生成对构建通用多模态交互AI系统至关重要。然而,扩散模型中双向注意力对所有视频帧同时去噪的迭代过程阻碍了实时交互。现有蒸馏方法虽能将模型变为自回归并减少采样步骤,但主要针对文本到视频生成,导致人机交互不自然且效率低。本文聚焦于基于多模态上下文(文本、图像、音频)的实时交互式视频扩散,以填补这一空白。针对主流的在线策略蒸馏方法Self Forcing在多模态条件下出现视觉伪影(如闪烁、黑帧、质量下降)的问题,我们提出改进蒸馏方案,重点关注条件输入质量以及在线优化的初始化与调度策略。在包含HDTF、AVSpeech、CelebV-HQ的多模态条件(音频、图像、文本)角色视频生成基准测试中,所提蒸馏模型在与同等或更大规模的全步双向基线相当的视觉质量下,推理成本和延迟降低20倍。进一步,将该模型与语音语言模型及长视频推理技术Anchor-Heavy Identity Sinks结合,构建了实时多模态交互角色系统LiveTalk。在自建的多轮交互基准上的系统级评估显示,LiveTalk在多轮视频连贯性和内容质量上优于Sora2、Veo3等先进模型,同时将响应延迟从1至2分钟缩短至实时生成,实现无缝人机多模态交互。
原文摘要 · Abstract (English)
Real-time video generation via diffusion is essential for building general-purpose multimodal interactive AI systems. However, the simultaneous denoising of all video frames with bidirectional attention via an iterative process in diffusion models prevents real-time interaction. While existing distillation methods can make the model autoregressive and reduce sampling steps to mitigate this, they focus primarily on text-to-video generation, leaving the human-AI interaction unnatural and less efficient. This paper targets real-time interactive video diffusion conditioned on a multimodal context, including text, image, and audio, to bridge the gap. Given the observation that the leading on-policy distillation approach Self Forcing encounters challenges (visual artifacts like flickering, black frames, and quality degradation) with multimodal conditioning, we investigate an improved distillation recipe with emphasis on the quality of condition inputs as well as the initialization and schedule for the on-policy optimization. On benchmarks for multimodal-conditioned (audio, image, and text) avatar video generation including HDTF, AVSpeech, and CelebV-HQ, our distilled model matches the visual quality of the full-step, bidirectional baselines of similar or larger size with 20x less inference cost and latency. Further, we integrate our model with audio language models and long-form video inference technique Anchor-Heavy Identity Sinks to build LiveTalk, a real-time multimodal interactive avatar system. System-level evaluation on our curated multi-turn interaction benchmark shows LiveTalk outperforms state-of-the-art models (Sora2, Veo3) in multi-turn video coherence and content quality, while reducing response latency from 1 to 2 minutes to real-time generation, enabling seamless human-AI multimodal interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。