提出新表征与评估标准,让数字人说话更自然逼真。
Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation Metrics
- 设计语音-网格同步表征,捕捉语音与嘴型的精细对应
- 引入三重评价指标,显著提升嘴型同步与表现力
- 可直接用于现有模型,提升生成质量且开源可用
近期语音驱动的3D人脸生成在唇部同步方面取得进展,但依然难以捕捉语音特征与嘴型运动之间的感知一致性。本文指出时序同步性、唇读可辨识性与表现力是实现感知准确唇动的三大关键。基于假设存在理想表示空间以满足这三项标准,我们提出一种语音-网格同步表征,能精确建模语音信号与3D人脸网格间的复杂对应关系。实验表明,该学习到的表示具有理想特性,将其作为感知损失融入现有模型,可有效提升唇动对齐效果。同时,我们利用该表示作为感知评估指标,并引入两个基于物理的唇同步度量,全面评估生成结果在三方面的表现。大量实验证明,使用该感知损失训练的模型在三个维度上均有显著提升。代码与数据集见 https://perceptual-3d-talking-head.github.io/。
原文摘要 · Abstract (English)
Recent advancements in speech-driven 3D talking head generation have made significant progress in lip synchronization. However, existing models still struggle to capture the perceptual alignment between varying speech characteristics and corresponding lip movements. In this work, we claim that three criteria -- Temporal Synchronization, Lip Readability, and Expressiveness -- are crucial for achieving perceptually accurate lip movements. Motivated by our hypothesis that a desirable representation space exists to meet these three criteria, we introduce a speech-mesh synchronized representation that captures intricate correspondences between speech signals and 3D face meshes. We found that our learned representation exhibits desirable characteristics, and we plug it into existing models as a perceptual loss to better align lip movements to the given speech. In addition, we utilize this representation as a perceptual metric and introduce two other physically grounded lip synchronization metrics to assess how well the generated 3D talking heads align with these three criteria. Experiments show that training 3D talking head generation models with our perceptual loss significantly improve all three aspects of perceptually accurate lip synchronization. Codes and datasets are available at https://perceptual-3d-talking-head.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。