用嘴部发音动作提升跨语料库语音情感识别效果
Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
- 以嘴部发音动作为核心,替代易变的声学特征
- 在CREMA-D和MSP-IMPROV数据集上验证了有效性
- 适合做跨场景语音情感分析的研究者参考
跨语料库语音情感识别(SER)在诸多实际应用中至关重要。传统方法多聚焦于对声学特征进行适应性调整以匹配不同语料库、领域或标签,但声学特征易受说话人差异、领域偏移和录音条件等因素影响,具有固有可变性和误差。为此,本研究提出一种新对比方法,将情绪特异性的发音动作作为分析核心。通过关注更稳定一致的发音动作,旨在提升SER任务中的情感迁移学习效果。研究以CREMA-D和MSP-IMPROV为基准数据集,揭示了这些发音动作的共性与可靠性。结果表明,嘴部发音动作具有更强约束力,能有效提升跨不同设置或领域的语音情感识别性能。
原文摘要 · Abstract (English)
Cross-corpus speech emotion recognition (SER) plays a vital role in numerous practical applications. Traditional approaches to cross-corpus emotion transfer often concentrate on adapting acoustic features to align with different corpora, domains, or labels. However, acoustic features are inherently variable and error-prone due to factors like speaker differences, domain shifts, and recording conditions. To address these challenges, this study adopts a novel contrastive approach by focusing on emotion-specific articulatory gestures as the core elements for analysis. By shifting the emphasis on the more stable and consistent articulatory gestures, we aim to enhance emotion transfer learning in SER tasks. Our research leverages the CREMA-D and MSP-IMPROV corpora as benchmarks and it reveals valuable insights into the commonality and reliability of these articulatory gestures. The findings highlight mouth articulatory gesture potential as a better constraint for improving emotion recognition across different settings or domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。