arXiv:2506.02239cs.CLeess.AS2025-06中稿 · Interspeech 2025被引 2

用词重要性筛选语音片段,提升情绪识别准确率

Investigating the Impact of Word Informativeness on Speech Emotion Recognition

  • 基于预训练模型计算词的重要程度,定位关键语音段
  • 仅在高信息量词段提取特征,准确率显著提升
  • 适合语音情绪分析、细粒度语音处理研究者

在语音情绪识别中,关键挑战在于识别携带最相关声学变化的语音片段。传统方法对整句或更长语音段计算能量、基频等功能统计量,可能遗漏细粒度变化。本研究利用预训练语言模型计算词的重要性,识别语义关键段,并仅在这些段落上计算声学特征(包括标准语调特征、其函数量及自监督表示),显著提升了情绪识别性能,验证了该方法的有效性。

原文摘要 · Abstract (English)

In emotion recognition from speech, a key challenge lies in identifying speech signal segments that carry the most relevant acoustic variations for discerning specific emotions. Traditional approaches compute functionals for features such as energy and F0 over entire sentences or longer speech portions, potentially missing essential fine-grained variation in the long-form statistics. This research investigates the use of word informativeness, derived from a pre-trained language model, to identify semantically important segments. Acoustic features are then computed exclusively for these identified segments, enhancing emotion recognition accuracy. The methodology utilizes standard acoustic prosodic features, their functionals, and self-supervised representations. Results indicate a notable improvement in recognition performance when features are computed on segments selected based on word informativeness, underscoring the effectiveness of this approach.

情绪识别语音分析语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。