通过声学轨迹建模,提升双人对话中情感同步的检测精度。
TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

- 将对话视为声学嵌入序列,捕捉时间关系
- 在新数据集上达97.01%准确率,优于传统方法
- 适合研究情感计算与交互式语音系统的学者
随着语音AI代理的普及,理解对话中情感同步的重要性日益凸显。情感同步受社交关系和对话上下文影响,随时间动态变化。我们构建了DyadEE数据集,包含真实情感同步对话及通过角色互换与情绪重合成生成的非同步对话。进一步提出TRACE框架,将双人互动建模为基于情感微调的Whisper声学嵌入序列,以窗口级方式处理每段交互痕迹,而非对话语句聚合。在DyadEE上的实验表明,引入对话上下文和关系信息可显著提升情感同步检测效果,TRACE达到97.01%最高准确率。
原文摘要 · Abstract (English)
With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly important. Emotional entrainment is shaped by social relationships and conversational context, influencing affective coordination over time. We introduce DyadEE, a dataset for emotional entrainment detection in dyadic speech interactions, containing both emotionally entrained conversations and synthetic interactions where entrainment is disrupted through partner swapping and emotion resynthesis. We further propose TRACE, a window-level framework that models dyadic interaction as ordered sequences of acoustic embeddings derived from emotion fine-tuned Whisper representations, treating each sample as an interaction trace rather than pooled utterances. Experimental results on DyadEE show that incorporating conversational context and relationship information improves emotional entrainment detection, with TRACE achieving the best accuracy of 97.01%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。