用视觉语音识别提升3D虚拟人唇动精度
VisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis
- 通过可微渲染与视觉语音模型联合训练,实现高保真唇动生成
- 在MEAD数据集上唇部误差降低56.1%,感知质量显著提升
- 适合需要精准口型的无障碍交互与手语虚拟人场景
逼真的高保真3D面部动画对人机交互和无障碍系统至关重要。尽管已有方法表现良好,但其依赖网格域,难以充分利用2D计算机视觉与图形学的快速进展。本文提出VisualSpeaker,利用光栅化渲染并结合预训练视觉自动语音识别模型,构建感知驱动的唇读损失函数,实现更精确的3D面部动画生成。在MEAD数据集上的评估显示,该方法使唇部顶点误差(Lip Vertex Error)下降56.1%,同时显著提升生成动画的感知质量,且保持了网格驱动动画的可控性。该感知导向机制自然支持准确的口型表达,为手语虚拟人中相似手势的歧义消解提供了关键线索。
原文摘要 · Abstract (English)
Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Although prior methods show promising quality, their reliance on the mesh domain limits their ability to fully leverage the rapid visual innovations seen in 2D computer vision and graphics. We propose VisualSpeaker, a novel method that bridges this gap using photorealistic differentiable rendering, supervised by visual speech recognition, for improved 3D facial animation. Our contribution is a perceptual lip-reading loss, derived by passing photorealistic 3D Gaussian Splatting avatar renders through a pre-trained Visual Automatic Speech Recognition model during training. Evaluation on the MEAD dataset demonstrates that VisualSpeaker improves both the standard Lip Vertex Error metric by 56.1% and the perceptual quality of the generated animations, while retaining the controllability of mesh-driven animation. This perceptual focus naturally supports accurate mouthings, essential cues that disambiguate similar manual signs in sign language avatars.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。