通过分粒度解耦与交替优化,提升语音驱动人脸生成的清晰度与稳定性。
M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation
- 分粒度解耦非刚性与刚性运动,独立建模面部与头部动作
- 在多个数据集上实现2.43 dB PSNR提升与0.64分真实感增益
- 适合影视制作中对高帧率、低伪影人脸视频的需求
语音驱动的人脸生成在影视制作中具有重要潜力。现有3D方法虽在运动建模与内容合成方面取得进展,但常因难以表示稳定精细的运动场而产生运动模糊、时间抖动和局部穿透等渲染伪影。通过系统分析,我们将人脸生成重构为统一框架:视频预处理、运动表征与渲染重建三步。基于此,提出M2DAO-Talker,通过多粒度运动解耦与交替优化解决现有问题。设计新型2D肖像预处理流程,提取逐帧形变控制条件(运动区域分割掩码与相机参数)以支持运动表征;提出多粒度运动解耦策略,独立建模非刚性(口部与面部)与刚性(头部)运动,提升重建精度;引入运动一致性约束,确保头-躯干运动连贯性,缓解由运动混叠引起的穿透伪影;同时设计交替优化策略,迭代精炼面部与口部运动参数,生成更真实视频。跨多个数据集实验表明,相比TalkingGaussian,M2DAO-Talker在生成质量上提升2.43 dB PSNR,用户评估视频真实感提升0.64分,推理速度达150 FPS。
原文摘要 · Abstract (English)
Audio-driven talking head generation holds significant potential for film production. While existing 3D methods have advanced motion modeling and content synthesis, they often produce rendering artifacts, such as motion blur, temporal jitter, and local penetration, due to limitations in representing stable, fine-grained motion fields. Through systematic analysis, we reformulate talking head generation into a unified framework comprising three steps: video preprocessing, motion representation, and rendering reconstruction. This framework underpins our proposed M2DAO-Talker, which addresses current limitations via multi-granular motion decoupling and alternating optimization. Specifically, we devise a novel 2D portrait preprocessing pipeline to extract frame-wise deformation control conditions (motion region segmentation masks, and camera parameters) to facilitate motion representation. To ameliorate motion modeling, we elaborate a multi-granular motion decoupling strategy, which independently models non-rigid (oral and facial) and rigid (head) motions for improved reconstruction accuracy. Meanwhile, a motion consistency constraint is developed to ensure head-torso kinematic consistency, thereby mitigating penetration artifacts caused by motion aliasing. In addition, an alternating optimization strategy is designed to iteratively refine facial and oral motion parameters, enabling more realistic video generation. Experiments across multiple datasets show that M2DAO-Talker achieves state-of-the-art performance, with the 2.43 dB PSNR improvement in generation quality and 0.64 gain in user-evaluated video realness versus TalkingGaussian while with 150 FPS inference speed. Our project homepage is https://m2dao-talker.github.io/M2DAO-Talk.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。