arXiv:2608.15296cs.CV2026-08中稿 · publication in IEE…被引 1

用真人偏好训练模型,让语音驱动的3D人脸动画更自然逼真。

FMReward: Aligning and Evaluating Audio-Driven 3D Facial Animation with Human Preferences

论文配图:FMReward: Aligning and Evaluating Audio-Driven 3D Facial Animation with Human Preferences
图 1 · 摘自论文原文
  • 基于真人对比数据构建感知评分模型,自动评估动画质量。
  • 在65,574对动画中验证,新模型比传统方法更贴近人类偏好。
  • 适合做虚拟人、游戏角色语音动画的研究者和开发者。

语音驱动的3D人脸动画对提升虚拟体验的沉浸感与互动性至关重要。尽管近期进展令人鼓舞,但现有方法的训练与评估多依赖于真实动作误差,难以匹配人类偏好。为此,我们提出一个完整框架:从真人偏好数据中学习感知模型,并用于提升与评估动画的主观质量。首先,我们构建了FMPair(面部运动成对偏好)数据集,这是首个面向语音驱动3D人脸动画的人类偏好数据集,通过系统化标注流程生成,包含来自8,834段真实音频片段的65,574对3D面部动作。基于该数据集,我们提出FMReward模型,输入音频与3D面部动作,输出与人类偏好一致的感知质量分数。在此基础上,我们进一步提出面部运动奖励反馈学习(FMFL),一种直接微调算法,利用预训练的奖励模型优化基于扩散模型的语音驱动3D人脸动画,使其更贴合人类感受。大量实验表明,FMReward在匹配人类偏好方面优于其他度量,而FMFL能显著提升动画的感知质量。

原文摘要 · Abstract (English)

Audio-driven 3D facial animation is essential for advancing immersion and interactivity in virtual experiences. Although recent advances have shown promising capabilities, the training and evaluation of existing methods typically rely on ground-truth-based errors, which fall short of aligning with human preferences. To address this, we present a comprehensive framework that learns an automatic perceptual model from human preference data and leverages it to improve and evaluate the perceptual quality of audio-driven 3D facial animation. To begin with, we construct FMPair (Facial Motion Pairwise preference), the first human preference dataset for audio-driven 3D facial animation, which is built through a systematic annotation pipeline and comprises 65,574 annotated 3D facial motion pairs from 8,834 distinct in-the-wild audio clips. Based on the pairwise comparison dataset, we propose a Facial Motion Reward model, termed FMReward, which takes audio and 3D facial motion as inputs and predicts a perceptual quality score aligned with human preferences. Building upon FMReward, we further introduce Facial Motion reward Feedback Learning (FMFL), a direct fine-tuning algorithm that leverages a pretrained reward model to optimize diffusion-based audio-driven 3D facial animation models for better alignment with human preferences. Extensive experiments demonstrate the superiority of FMReward over other metrics in aligning with human preferences and the effectiveness of FMFL in improving the perceptual quality of audio-driven 3D facial animation.

3D人脸动画语音驱动偏好学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。