用扩散模型加速对话中听众面部动作生成,实时性提升99%。
Efficient Listener: Dyadic Facial Motion Synthesis via Action Diffusion
- 引入图像生成领域的扩散模型生成面部动作,突破3DMM计算瓶颈。
- 在多个数据集上实现比当前最佳方法更高的真实感与同步性。
- 适合需要低延迟交互的虚拟助手、数字人等应用场景。
在双人对话中生成逼真的听众面部动作仍具挑战,源于动作空间高维与时间依赖性强。现有方法通常提取3D形态模型(3DMM)系数并在其空间建模,但3DMM计算效率低下,难以实现实时交互响应。为此,我们提出面部动作扩散(FAD),将图像生成领域的扩散方法引入面部动作生成。同时构建专用的高效听众网络(ELNet),融合说话者视觉与音频信息作为输入。结合FAD与ELNet,所提方法学习有效的听众面部动作表示,在性能超越当前最优方法的同时,计算时间减少99%。
原文摘要 · Abstract (English)
Generating realistic listener facial motions in dyadic conversations remains challenging due to the high-dimensional action space and temporal dependency requirements. Existing approaches usually consider extracting 3D Morphable Model (3DMM) coefficients and modeling in the 3DMM space. However, this makes the computational speed of the 3DMM a bottleneck, making it difficult to achieve real-time interactive responses. To tackle this problem, we propose Facial Action Diffusion (FAD), which introduces the diffusion methods from the field of image generation to achieve efficient facial action generation. We further build the Efficient Listener Network (ELNet) specially designed to accommodate both the visual and audio information of the speaker as input. Considering of FAD and ELNet, the proposed method learns effective listener facial motion representations and leads to improvements of performance over the state-of-the-art methods while reducing 99% computational time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。