通过语音与说话人风格对齐,提升跨语言情感识别效果
Speaker Style-Aware Phoneme Anchoring for Improved Cross-Lingual Speech Emotion Recognition
- 构建基于图聚类的说话人情感社区,捕捉共性表达特征
- 在语音与说话人双空间中锚定情感,跨语言识别准确率提升
- 适用于多语言情感分析、个性化语音交互场景
跨语言语音情感识别(SER)因发音差异和说话人特异性表达风格不同而面临挑战。为有效捕捉跨语言的情感表达,本文提出一种说话人风格感知的音素锚定框架,实现语音与说话人层面的情感对齐。通过图聚类构建情感特定的说话人社群,以捕捉共享的说话人特征;在此基础上,在说话人空间与音素空间中实施双空间锚定,促进情感跨语言迁移。在MSP-Podcast(英语)与BIIC-Podcast(台湾国语)数据集上的评估显示,该方法在多个基线模型上实现了更好的泛化性能,并揭示了跨语言情感表征中的共性规律。
原文摘要 · Abstract (English)
Cross-lingual speech emotion recognition (SER) remains a challenging task due to differences in phonetic variability and speaker-specific expressive styles across languages. Effectively capturing emotion under such diverse conditions requires a framework that can align the externalization of emotions across different speakers and languages. To address this problem, we propose a speaker-style aware phoneme anchoring framework that aligns emotional expression at the phonetic and speaker levels. Our method builds emotion-specific speaker communities via graph-based clustering to capture shared speaker traits. Using these groups, we apply dual-space anchoring in speaker and phonetic spaces to enable better emotion transfer across languages. Evaluations on the MSP-Podcast (English) and BIIC-Podcast (Taiwanese Mandarin) corpora demonstrate improved generalization over competitive baselines and provide valuable insights into the commonalities in cross-lingual emotion representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。