通过关键帧增强的双路径扩散模型,提升语音驱动人脸动画的精准与自然。
KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation
- 分离语音中的表情与头部姿态特征,实现细粒度控制。
- 预测关键动态帧,显著提升唇同步与头部动作自然度。
- 适合需要高精度人脸动画的视频生成与虚拟角色应用。
语音驱动的人脸动画在多媒体应用中取得显著进展,扩散模型在说话人脸合成方面展现出强大潜力。然而,现有方法通常将语音特征视为单一整体表示,未能捕捉其对不同面部运动的精细驱动作用,同时忽略了高强度动态关键帧建模的重要性。为此,我们提出KSDiff,一种关键帧增强的语音感知双路径扩散框架。具体而言,原始音频和转录文本通过双路径语音编码器(DPSE)解耦表达相关与头部姿态相关特征,同时采用自回归关键帧建立学习(KEL)模块预测最具表现力的运动帧。这些组件被整合至双路径运动生成器中,以合成连贯且逼真的面部动作。在HDTF和VoxCeleb数据集上的大量实验表明,KSDiff达到最先进性能,唇同步准确率与头部姿态自然度均有提升。结果验证了结合语音解耦与关键帧感知扩散在说话头生成中的有效性。演示页面见:https://kincin.github.io/KSDiff/。
原文摘要 · Abstract (English)
Audio-driven facial animation has made significant progress in multimedia applications, with diffusion models showing strong potential for talking-face synthesis. However, most existing works treat speech features as a monolithic representation and fail to capture their fine-grained roles in driving different facial motions, while also overlooking the importance of modeling keyframes with intense dynamics. To address these limitations, we propose KSDiff, a Keyframe-Augmented Speech-Aware Dual-Path Diffusion framework. Specifically, the raw audio and transcript are processed by a Dual-Path Speech Encoder (DPSE) to disentangle expression-related and head-pose-related features, while an autoregressive Keyframe Establishment Learning (KEL) module predicts the most salient motion frames. These components are integrated into a Dual-path Motion generator to synthesize coherent and realistic facial motions. Extensive experiments on HDTF and VoxCeleb demonstrate that KSDiff achieves state-of-the-art performance, with improvements in both lip synchronization accuracy and head-pose naturalness. Our results highlight the effectiveness of combining speech disentanglement with keyframe-aware diffusion for talking-head generation. The demo page is available at: https://kincin.github.io/KSDiff/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。