实时驱动人脸动画,精准同步语音与口型。
RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
- 采用混合注意力机制实现音频细粒度控制
- 在实时约束下保持高画质与口型同步
- 无需显式运动监督,避免长期时序漂移
音频驱动的人脸动画旨在从输入音频和单张参考图像生成逼真自然的说话头视频。现有方法虽通过高维中间表示和显式建模运动动态实现高质量结果,但计算复杂度高,难以实现实时部署。实时推理对延迟和内存有严格要求,常需使用高度压缩的潜在表示,但这会损害细微时空细节的保留,导致音画不同步。本文提出RAP(Real-time Audio-driven Portrait animation),一个在实时约束下生成高质量说话人脸的统一框架。具体而言,RAP引入混合注意力机制以实现细粒度音频控制,并采用静态-动态训练-推理范式,避免显式运动监督。通过这些技术,RAP实现了精确的音频驱动控制,缓解了长期时序漂移,同时保持了高视觉保真度。大量实验表明,RAP在实时约束下达到当前最佳性能。
原文摘要 · Abstract (English)
Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamics, their computational complexity renders them unsuitable for real-time deployment. Real-time inference imposes stringent latency and memory constraints, often necessitating the use of highly compressed latent representations. However, operating in such compact spaces hinders the preservation of fine-grained spatiotemporal details, thereby complicating audio-visual synchronization RAP (Real-time Audio-driven Portrait animation), a unified framework for generating high-quality talking portraits under real-time constraints. Specifically, RAP introduces a hybrid attention mechanism for fine-grained audio control, and a static-dynamic training-inference paradigm that avoids explicit motion supervision. Through these techniques, RAP achieves precise audio-driven control, mitigates long-term temporal drift, and maintains high visual fidelity. Extensive experiments demonstrate that RAP achieves state-of-the-art performance while operating under real-time constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。