arXiv:2505.13455eess.AScs.AI2025-05被引 1

研究对话中语音重叠如何影响面部与声音情绪同步

Spatiotemporal Emotional Synchrony in Dyadic Interactions: The Role of Speech Conditions in Facial and Vocal Affective Alignment

  • 对比非重叠与重叠语音下的表情与语音情绪同步性
  • 非重叠语音下情绪同步更稳定,尤其在唤醒度上变异小
  • 面部常先于语音,语音则主导同时发声时的情绪引导

理解人类在多通道交流中(尤其是面部表情与语音)情绪表达与同步的机制,对情感识别系统和人机交互具有重要意义。基于非重叠语音促进情绪协调、重叠语音破坏同步的假设,本研究利用IEMOCAP数据集中的二人互动片段,通过EmoNet(面部视频)和基于Wav2Vec2的模型(语音音频)提取连续情绪估计值。根据语音重叠情况分类,采用皮尔逊相关、时移校正分析及动态时间规整(DTW)评估情绪对齐。结果表明:非重叠语音段情绪同步更稳定可预测,尽管零时滞相关性低且无统计差异,但其变异性更低,尤其在唤醒度维度;时移相关与最佳时滞分布显示非重叠段存在更清晰一致的时间对齐。相反,重叠语音段变异性更高、时滞分布平坦,但DTW揭示其仍具紧密对齐,暗示不同协调策略。方向性分析发现,换轮次时面部表情常领先语音,而同时发声时语音主导。研究强调对话结构在调节情绪交流中的关键作用,为真实交互中多模态情感对齐的时空动态提供了新见解。

原文摘要 · Abstract (English)

Understanding how humans express and synchronize emotions across multiple communication channels particularly facial expressions and speech has significant implications for emotion recognition systems and human computer interaction. Motivated by the notion that non-overlapping speech promotes clearer emotional coordination, while overlapping speech disrupts synchrony, this study examines how these conversational dynamics shape the spatial and temporal alignment of arousal and valence across facial and vocal modalities. Using dyadic interactions from the IEMOCAP dataset, we extracted continuous emotion estimates via EmoNet (facial video) and a Wav2Vec2-based model (speech audio). Segments were categorized based on speech overlap, and emotional alignment was assessed using Pearson correlation, lag adjusted analysis, and Dynamic Time Warping (DTW). Across analyses, non overlapping speech was associated with more stable and predictable emotional synchrony than overlapping speech. While zero-lag correlations were low and not statistically different, non overlapping speech showed reduced variability, especially for arousal. Lag adjusted correlations and best-lag distributions revealed clearer, more consistent temporal alignment in these segments. In contrast, overlapping speech exhibited higher variability and flatter lag profiles, though DTW indicated unexpectedly tighter alignment suggesting distinct coordination strategies. Notably, directionality patterns showed that facial expressions more often preceded speech during turn-taking, while speech led during simultaneous vocalizations. These findings underscore the importance of conversational structure in regulating emotional communication and provide new insight into the spatial and temporal dynamics of multimodal affective alignment in real world interaction.

情绪同步多模态语音交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。