用高斯点云实现更真实的单图说话头生成
Splat-Portrait: Generalizing Talking Heads with Gaussian Splatting
- 将单张人脸图分解为静态3D点云和2D背景,自动分离
- 仅凭音频生成自然唇动,无需运动先验或3D标注
- 无3D监督训练,视觉质量优于现有方法
说话头生成旨在从语音和单张肖像图合成自然的动态视频。以往基于3D的方法依赖领域特定的启发式规则(如基于形变的面部运动表示先验),仍导致3D人物重建不准确,影响动画真实感。本文提出Splat-Portrait,一种基于高斯点云的方法,解决3D头部重建与唇部运动合成难题。该方法自动将单张肖像图解耦为静态3D重建(以静态高斯点云表示)和全图2D背景;并根据输入音频生成自然唇动,无需任何运动驱动先验。训练仅依赖2D重建损失与得分蒸馏损失,无需3D监督或关键点标注。实验表明,Splat-Portrait在说话头生成与新视角合成任务上表现优异,视觉质量超越先前工作。
原文摘要 · Abstract (English)
Talking Head Generation aims at synthesizing natural-looking talking videos from speech and a single portrait image. Previous 3D talking head generation methods have relied on domain-specific heuristics such as warping-based facial motion representation priors to animate talking motions, yet still produce inaccurate 3D avatar reconstructions, thus undermining the realism of generated animations. We introduce Splat-Portrait, a Gaussian-splatting-based method that addresses the challenges of 3D head reconstruction and lip motion synthesis. Our approach automatically learns to disentangle a single portrait image into a static 3D reconstruction represented as static Gaussian Splatting, and a predicted whole-image 2D background. It then generates natural lip motion conditioned on input audio, without any motion driven priors. Training is driven purely by 2D reconstruction and score-distillation losses, without 3D supervision nor landmarks. Experimental results demonstrate that Splat-Portrait exhibits superior performance on talking head generation and novel view synthesis, achieving better visual quality compared to previous works. Our project code and supplementary documents are public available at https://github.com/stonewalking/Splat-portrait.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。