arXiv:2512.16081cs.HCcs.AI2025-12被引 7

在虚拟现实中评估生成模型的情感3D动画,发现情感建模比单纯同步语音更真实可信。

Evaluation of Generative Models for Emotional 3D Animation Generation in VR

  • 在VR中通过用户感知指标评估三种语音驱动的3D动画方法
  • 情感建模方法识别准确率更高,快乐表情比中性表情更自然真实
  • 生成模型在面部表现和互动体验上仍落后于真人重建,但多样性受好评

社交互动中,非语言信号如面部表情和身体动作与言语共同传递情绪。生成模型在同步语音生成全身非语言动画方面已取得良好进展;然而,基于2D场景的统计指标无法充分捕捉用户感知的情绪,限制了对模型有效性的理解。为此,我们在虚拟现实(VR)环境中评估情感3D动画生成模型,重点考察用户中心指标:情绪唤醒度、真实性、自然性、愉悦感、多样性及实时人机交互质量。通过48名参与者的人机交互实验,我们评估了三种前沿语音驱动3D动画方法在“快乐”(高唤醒)和“中性”(中等唤醒)两种情绪下的感知情绪质量,并与基于重建的方法生成的真实人类表情进行对比,以分析其优劣及对真实表情的还原程度。结果表明,显式建模情绪的方法相比仅关注语音同步的方法具有更高的情绪识别准确率。用户对快乐动画的真实性与自然性评价显著高于中性动画,凸显当前模型在处理细微情绪状态时的局限性。生成模型在面部表情质量上不及重建方法,且所有方法在动画愉悦感和交互质量上得分均较低,强调了将用户中心评估纳入生成模型开发的重要性。此外,参与者普遍认可各生成模型在动画多样性方面的表现。

原文摘要 · Abstract (English)

Social interactions incorporate nonverbal signals to convey emotions alongside speech, including facial expressions and body gestures. Generative models have demonstrated promising results in creating full-body nonverbal animations synchronized with speech; however, evaluations using statistical metrics in 2D settings fail to fully capture user-perceived emotions, limiting our understanding of model effectiveness. To address this, we evaluate emotional 3D animation generative models within a Virtual Reality (VR) environment, emphasizing user-centric metrics emotional arousal realism, naturalness, enjoyment, diversity, and interaction quality in a real-time human-agent interaction scenario. Through a user study (N=48), we examine perceived emotional quality for three state of the art speech-driven 3D animation methods across two emotions happiness (high arousal) and neutral (mid arousal). Additionally, we compare these generative models against real human expressions obtained via a reconstruction-based method to assess both their strengths and limitations and how closely they replicate real human facial and body expressions. Our results demonstrate that methods explicitly modeling emotions lead to higher recognition accuracy compared to those focusing solely on speech-driven synchrony. Users rated the realism and naturalness of happy animations significantly higher than those of neutral animations, highlighting the limitations of current generative models in handling subtle emotional states. Generative models underperformed compared to reconstruction-based methods in facial expression quality, and all methods received relatively low ratings for animation enjoyment and interaction quality, emphasizing the importance of incorporating user-centric evaluations into generative model development. Finally, participants positively recognized animation diversity across all generative models.

3D动画情感生成虚拟现实用户评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。