用人类反馈提升动态语音情感识别,让虚拟人表情更自然。
Human Feedback Driven Dynamic Speech Emotion Recognition
- 分阶段训练模型,融合真实情感数据与合成情感序列。
- 基于狄利克雷分布建模情绪混合,比滑动窗口法更准确。
- 人类反馈可优化模型且简化标注,适合虚拟角色情感生成。
本文探索动态语音情感识别新方向,假设每段音频在不同时间点对应一系列持续变化的情绪。研究聚焦于3D虚拟角色的表情动画,提出多阶段方法:先训练传统语音情感识别模型,再合成情感序列,最后基于人类反馈进行模型优化。此外,引入狄利克雷分布建模情绪混合,提升表达能力。模型在3D面部动画数据集上评估,对比滑动窗口法,结果表明狄利克雷方法能更有效建模情绪混合。结合人类反馈进一步提升模型质量,同时简化标注流程。
原文摘要 · Abstract (English)
This work proposes to explore a new area of dynamic speech emotion recognition. Unlike traditional methods, we assume that each audio track is associated with a sequence of emotions active at different moments in time. The study particularly focuses on the animation of emotional 3D avatars. We propose a multi-stage method that includes the training of a classical speech emotion recognition model, synthetic generation of emotional sequences, and further model improvement based on human feedback. Additionally, we introduce a novel approach to modeling emotional mixtures based on the Dirichlet distribution. The models are evaluated based on ground-truth emotions extracted from a dataset of 3D facial animations. We compare our models against the sliding window approach. Our experimental results show the effectiveness of Dirichlet-based approach in modeling emotional mixtures. Incorporating human feedback further improves the model quality while providing a simplified annotation procedure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。