arXiv:2507.19356cs.CLcs.SD2025-07

对齐语音转录与说话人分离的时间戳,提升对话情感识别准确率。

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

  • 用预训练模型同步语音转录与说话人分离的时间戳,生成精准发言段。
  • 在IEMOCAP数据集上,对齐后情感识别准确率显著优于未对齐基线。
  • 适合做多模态情感分析、对话系统优化的研究者参考。

本文研究将自动语音识别(ASR)转录文本与说话人分离(SD)输出进行基于时间戳对齐,对语音情感识别(SER)准确率的影响。在对话场景中,两者间的时间错位常降低多模态情感识别系统的可靠性。为此,我们提出一个对齐流程,利用预训练的ASR与说话人分离模型,系统性地同步时间戳,生成精确标注的说话人片段。所提多模态方法结合RoBERTa提取的文本嵌入与Wav2Vec提取的音频嵌入,通过带门控机制的交叉注意力融合。在IEMOCAP基准数据集上的实验表明,精确的时间对齐提升了SER准确率,优于缺乏同步的基线方法。结果凸显了时间对齐的关键作用,证明其在增强整体情感识别性能方面的有效性,并为鲁棒的多模态情感分析提供了基础。

原文摘要 · Abstract (English)

In this paper, we investigate the impact of incorporating timestamp-based alignment between Automatic Speech Recognition (ASR) transcripts and Speaker Diarization (SD) outputs on Speech Emotion Recognition (SER) accuracy. Misalignment between these two modalities often reduces the reliability of multimodal emotion recognition systems, particularly in conversational contexts. To address this issue, we introduce an alignment pipeline utilizing pre-trained ASR and speaker diarization models, systematically synchronizing timestamps to generate accurately labeled speaker segments. Our multimodal approach combines textual embeddings extracted via RoBERTa with audio embeddings from Wav2Vec, leveraging cross-attention fusion enhanced by a gating mechanism. Experimental evaluations on the IEMOCAP benchmark dataset demonstrate that precise timestamp alignment improves SER accuracy, outperforming baseline methods that lack synchronization. The results highlight the critical importance of temporal alignment, demonstrating its effectiveness in enhancing overall emotion recognition accuracy and providing a foundation for robust multimodal emotion analysis.

情感识别语音处理多模态时间对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。