研究对话中语音如何触发面部和手势情绪表达,揭示互动时序对情感同步的影响。
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
- 基于双人交互数据,分析语音与面部/手部动作的区域特异性耦合关系。
- 非重叠语音引发更强面部活动,悲伤更明显,愤怒在重叠时抑制手势。
- 语音到手势预测中,语调和梅尔倒谱系数效果最佳,适合情感计算系统优化。
人类情绪表达通过语音、面部和手势信号协同产生。尽管语音与面部对齐已较为成熟,但情绪化语音与面部及手部运动之间的动态关联仍需深入理解,以揭示真实互动中情绪与行为线索的传递机制。对话结构如轮流发言形成稳定的多模态同步窗口,而同时说话(常代表高唤醒状态)则破坏对齐,影响情绪清晰度。本研究利用IEMOCAP数据集中的双人交互三维动作捕捉数据,分析语音特征(低层级语调、MFCC、模型推断的唤醒度、效价及类别情绪:快乐、悲伤、愤怒、中性)与面部及手部标记位移的耦合关系。通过帧级位移幅度量化表达活跃度,建立语音特征到面部和手部运动的预测映射。结果显示,非重叠语音显著提升下颌及口部活跃度;悲伤在非重叠时表达更强烈,愤怒在重叠时抑制手势。预测性能方面,语调和MFCC在发音区域预测准确率最高,而唤醒度和效价相关性较低且受上下文影响大。值得注意的是,手部与语音同步在低唤醒度和重叠语音下增强,但效价无此效应。
原文摘要 · Abstract (English)
Human emotional expression emerges through coordinated vocal, facial, and gestural signals. While speech face alignment is well established, the broader dynamics linking emotionally expressive speech to regional facial and hand motion remains critical for gaining a deeper insight into how emotional and behavior cues are communicated in real interactions. Further modulating the coordination is the structure of conversational exchange like sequential turn taking, which creates stable temporal windows for multimodal synchrony, and simultaneous speech, often indicative of high arousal moments, disrupts this alignment and impacts emotional clarity. Understanding these dynamics enhances realtime emotion detection by improving the accuracy of timing and synchrony across modalities in both human interactions and AI systems. This study examines multimodal emotion coupling using region specific motion capture from dyadic interactions in the IEMOCAP corpus. Speech features included low level prosody, MFCCs, and model derived arousal, valence, and categorical emotions (Happy, Sad, Angry, Neutral), aligned with 3D facial and hand marker displacements. Expressive activeness was quantified through framewise displacement magnitudes, and speech to gesture prediction mapped speech features to facial and hand movements. Nonoverlapping speech consistently elicited greater activeness particularly in the lower face and mouth. Sadness showed increased expressivity during nonoverlap, while anger suppressed gestures during overlaps. Predictive mapping revealed highest accuracy for prosody and MFCCs in articulatory regions while arousal and valence had lower and more context sensitive correlations. Notably, hand speech synchrony was enhanced under low arousal and overlapping speech, but not for valence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。