突破固定拓扑限制,实现任意3D人脸网格的高保真语音驱动动画
Beyond Fixed Topologies: Unregistered Training and Comprehensive Evaluation Metrics for 3D Talking Heads
- 用热扩散预测对拓扑不敏感的特征,支持任意网格结构动画
- 提出注册与非注册两种训练方式,首次实现真实扫描数据的动画生成
- 新评估指标更精准衡量唇同步效果,适合真实场景应用
语音驱动3D人脸动画面临网格拓扑不一致的挑战,现有方法依赖固定拓扑结构。本文首次提出可处理任意拓扑结构的框架,包括真实扫描数据。通过热扩散机制预测对拓扑鲁棒的特征,探索注册与完全非注册两种训练方式。实验表明,该方法在性能上优于固定拓扑技术,显著提升动画质量。同时,针对现有评估指标不足,提出新的唇同步评价指标。大量实验证明,该方案在灵活性和保真度上均达到新基准。代码与预训练模型已开源。
原文摘要 · Abstract (English)
Generating speech-driven 3D talking heads presents numerous challenges; among those is dealing with varying mesh topologies where no point-wise correspondence exists across the meshes the model can animate. While previous literature works assume fixed mesh structures, in this work we present the first framework capable of animating 3D faces in arbitrary topologies, including real scanned data. Our approach leverages heat diffusion to predict features that are robust to the mesh topology. We explore two training settings: a registered one, in which meshes in a training sequences share a fixed topology but any mesh can be animated at test time, and an fully unregistered one, which allows effective training with varying mesh structures. Additionally, we highlight the limitations of current evaluation metrics and propose new metrics for better lip-syncing evaluation. An extensive evaluation shows our approach performs favorably compared to fixed topology techniques, setting a new benchmark by offering a versatile and high-fidelity solution for 3D talking heads where the topology constraint is dropped. The code along with the pre-trained model are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。