让口型同步更准,表情控制更细,实时生成高保真人脸动画。
Takin-ADA: Emotion Controllable Audio-Driven Animation with Canonical and Landmark Loss Optimization
- 分两阶段优化:先增强细微表情传递,再提升口型同步精度。
- 512x512分辨率下达42帧/秒,性能超越现有商业方案。
- 支持灵活的表情与头部动作控制,适合影视级虚拟人应用。
现有音频驱动人脸动画方法存在表情泄露、细微表情传递无效及音画不同步等问题。我们发现根源在于运动表征不足和表情控制粒度不够。为此提出Takin-ADA,一种两阶段实时人脸动画生成方法。第一阶段引入专用损失函数,强化细微表情传递并减少无关表情泄露;第二阶段采用先进音频处理技术,提升唇形同步精度。该方法不仅能生成精准的口型动作,还支持灵活的表情与头部运动控制。在RTX 4090显卡上实现最高42 FPS的512x512高清动画生成,优于现有商业解决方案。大量实验表明,模型在视频质量、面部动态真实感和自然头部动作方面显著超越此前方法,树立了音频驱动人脸动画的新基准。
原文摘要 · Abstract (English)
Existing audio-driven facial animation methods face critical challenges, including expression leakage, ineffective subtle expression transfer, and imprecise audio-driven synchronization. We discovered that these issues stem from limitations in motion representation and the lack of fine-grained control over facial expressions. To address these problems, we present Takin-ADA, a novel two-stage approach for real-time audio-driven portrait animation. In the first stage, we introduce a specialized loss function that enhances subtle expression transfer while reducing unwanted expression leakage. The second stage utilizes an advanced audio processing technique to improve lip-sync accuracy. Our method not only generates precise lip movements but also allows flexible control over facial expressions and head motions. Takin-ADA achieves high-resolution (512x512) facial animations at up to 42 FPS on an RTX 4090 GPU, outperforming existing commercial solutions. Extensive experiments demonstrate that our model significantly surpasses previous methods in video quality, facial dynamics realism, and natural head movements, setting a new benchmark in the field of audio-driven facial animation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。