用高斯点云实现高保真、高速度的说话头像同步生成
SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting
- 基于高斯点云动态渲染,保持身份一致性
- 实现101帧/秒速度,唇动与语音同步率显著提升
- 适合需要实时高真实感视频生成的场景
生成逼真且语音驱动的说话头像视频面临严重同步挑战。一个自然的说话头像需协调身份特征、口型动作、面部表情和头部姿态。缺乏这些同步会导致结果失真。针对这一关键问题,我们提出SyncTalk++,采用基于高斯点云的动态肖像渲染器以保持身份一致,并设计面部同步控制器,通过3D面部混合形状模型精准重建表情,实现口型与语音同步。为提升头部运动自然性,引入头部同步稳定器优化姿态稳定性。此外,通过表情生成器和躯干恢复模块增强对分布外音频的鲁棒性,生成匹配的面部表情和无缝躯干区域。该方法在帧间保持视觉细节连续性,显著提升渲染速度与质量,最高达101帧每秒。大量实验与用户评估表明,SyncTalk++在同步性和真实感方面优于现有最优方法。
原文摘要 · Abstract (English)
Achieving high synchronization in the synthesis of realistic, speech-driven talking head videos presents a significant challenge. A lifelike talking head requires synchronized coordination of subject identity, lip movements, facial expressions, and head poses. The absence of these synchronizations is a fundamental flaw, leading to unrealistic results. To address the critical issue of synchronization, identified as the ''devil'' in creating realistic talking heads, we introduce SyncTalk++, which features a Dynamic Portrait Renderer with Gaussian Splatting to ensure consistent subject identity preservation and a Face-Sync Controller that aligns lip movements with speech while innovatively using a 3D facial blendshape model to reconstruct accurate facial expressions. To ensure natural head movements, we propose a Head-Sync Stabilizer, which optimizes head poses for greater stability. Additionally, SyncTalk++ enhances robustness to out-of-distribution (OOD) audio by incorporating an Expression Generator and a Torso Restorer, which generate speech-matched facial expressions and seamless torso regions. Our approach maintains consistency and continuity in visual details across frames and significantly improves rendering speed and quality, achieving up to 101 frames per second. Extensive experiments and user studies demonstrate that SyncTalk++ outperforms state-of-the-art methods in synchronization and realism. We recommend watching the supplementary video: https://ziqiaopeng.github.io/synctalk++.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。