用奖励模型从自动排序演示中生成更富情感的3D说话动画
ESARM: 3D Emotional Speech-to-Animation via Reward Model from Automatically-Ranked Demonstrations
- 通过交叉耦合训练解耦语音中的情绪与内容
- 自动评估生成动画质量,提升多样性与情感表现力
- 适合需要高情感真实性的虚拟角色动画开发
本文提出一种新型3D语音到动画(STA)生成框架,旨在解决现有模型在生成动画时缺乏情感深度和多样性的问题。当前的STA模型常生成不符合人类预期的平淡动画。为此,我们引入一个结合奖励模型的新颖STA模型,通过交叉耦合训练方法,在音频条件下实现情绪与内容的解耦。同时,设计了一种利用自动生成的面部动画质量评估来指导强化学习的训练策略,促使模型探索更广泛的可能性,从而生成高质量、多样化且情感丰富的3D面部动画。我们在基准数据集上进行了大量实验,结果验证了该框架在生成符合人类偏好的高质量、情感丰富3D动画方面的有效性。
原文摘要 · Abstract (English)
This paper proposes a novel 3D speech-to-animation (STA) generation framework designed to address the shortcomings of existing models in producing diverse and emotionally resonant animations. Current STA models often generate animations that lack emotional depth and variety, failing to align with human expectations. To overcome these limitations, we introduce a novel STA model coupled with a reward model. This combination enables the decoupling of emotion and content under audio conditions through a cross-coupling training approach. Additionally, we develop a training methodology that leverages automatic quality evaluation of generated facial animations to guide the reinforcement learning process. This methodology encourages the STA model to explore a broader range of possibilities, resulting in the generation of diverse and emotionally expressive facial animations of superior quality. We conduct extensive empirical experiments on a benchmark dataset, and the results validate the effectiveness of our proposed framework in generating high-quality, emotionally rich 3D animations that are better aligned with human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。