用语音生成逼真儿童虚拟人表情,提升心理训练沉浸感
Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications
- 结合Unreal Engine与Audio2Face实现实时情感面部生成
- 声音与表情不匹配会降低真实感,尤其愤怒情绪识别差
- 静音反而提升感知真实度,凸显音视频同步重要性
动态面部表情对生成可信的AI虚拟人至关重要,但多数系统仍为静态,限制了在虐待儿童调查模拟等虚拟训练中的应用。本文提出一种实时架构,结合Unreal Engine 5 MetaHuman渲染与NVIDIA Omniverse Audio2Face,基于语音韵律生成逼真儿童虚拟人的面部表情。由于文本转语音选项有限,两个虚拟人采用年轻女性成年模型配音,导致声音年龄与角色不符,可能影响视听一致性。采用双机配置分离语音生成与高负载渲染,实现桌面与VR端低延迟交互。通过70名参与者(被试间设计)比较音频+视觉与仅视觉条件,评估其对喜悦、悲伤和愤怒情绪的清晰度、面部真实感及共情程度的评分。结果表明,情绪总体可识别,尤以悲伤与喜悦为佳,但愤怒在无音频时更难辨识,凸显声音对高唤醒情绪的关键作用。有趣的是,静音片段反而提升了感知真实感,尤其当声调或年龄不匹配时。研究强调视听一致性的重要性:声音与动画不匹配会削弱表达效果,而良好匹配可弥补视觉不足,这对敏感场景中情感连贯虚拟人的构建提出挑战。
原文摘要 · Abstract (English)
Dynamic facial emotion is essential for believable AI-generated avatars, yet most systems remain visually static, limiting their use in simulations like virtual training for investigative interviews with abused children. We present a real-time architecture combining Unreal Engine 5 MetaHuman rendering with NVIDIA Omniverse Audio2Face to generate facial expressions from vocal prosody in photorealistic child avatars. Due to limited TTS options, both avatars were voiced using young adult female models from two systems to better fit character profiles, introducing a voice-age mismatch. This confound may affect audiovisual alignment. We used a two-PC setup to decouple speech generation from GPU-intensive rendering, enabling low-latency interaction in desktop and VR. A between-subjects study (N=70) compared audio+visual vs. visual-only conditions as participants rated emotional clarity, facial realism, and empathy for avatars expressing joy, sadness, and anger. While emotions were generally recognized - especially sadness and joy - anger was harder to detect without audio, highlighting the role of voice in high-arousal expressions. Interestingly, silencing clips improved perceived realism by removing mismatches between voice and animation, especially when tone or age felt incongruent. These results emphasize the importance of audiovisual congruence: mismatched voice undermines expression, while a good match can enhance weaker visuals - posing challenges for emotionally coherent avatars in sensitive contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。