用语音实时驱动逼真3D人脸动画,延迟低于15毫秒。
Audio Driven Real-Time Facial Animation for Social Telepresence
- 音频转表情序列的编码器+扩散模型解码,实现低延迟生成。
- 单步去噪加速,推理速度比现有方法快100到1000倍。
- 适用于虚拟社交、多语言演讲等实时场景,支持情绪与眼动输入。
我们提出一种面向虚拟现实社交互动的音频驱动实时3D人脸动画系统,可实现极低延迟(<15ms GPU时间)的逼真人脸生成。核心是将音频信号实时转换为面部表情潜空间序列,并通过扩散模型解码生成高质量3D头像。创新性地采用在线Transformer消除未来信息依赖,结合蒸馏流水线将迭代去噪压缩为单步,显著降低延迟。系统在连续音频帧处理中保持动画一致性,支持情绪条件与头戴式眼动传感器等多模态输入。实验表明,相比现有离线最优方法,动画精度显著提升,推理速度提升100至1000倍。通过实时VR演示验证了在多语言演讲等多种场景下的有效性。
原文摘要 · Abstract (English)
We present an audio-driven real-time system for animating photorealistic 3D facial avatars with minimal latency, designed for social interactions in virtual reality for anyone. Central to our approach is an encoder model that transforms audio signals into latent facial expression sequences in real time, which are then decoded as photorealistic 3D facial avatars. Leveraging the generative capabilities of diffusion models, we capture the rich spectrum of facial expressions necessary for natural communication while achieving real-time performance (<15ms GPU time). Our novel architecture minimizes latency through two key innovations: an online transformer that eliminates dependency on future inputs and a distillation pipeline that accelerates iterative denoising into a single step. We further address critical design challenges in live scenarios for processing continuous audio signals frame-by-frame while maintaining consistent animation quality. The versatility of our framework extends to multimodal applications, including semantic modalities such as emotion conditions and multimodal sensors with head-mounted eye cameras on VR headsets. Experimental results demonstrate significant improvements in facial animation accuracy over existing offline state-of-the-art baselines, achieving 100 to 1000 times faster inference speed. We validate our approach through live VR demonstrations and across various scenarios such as multilingual speeches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。