构建多语言多情感语音强调数据集,验证模型跨语言情感泛化能力
Do Speech Emphasis Models Generalize across Languages and Emotions?

- 构建涵盖7语言34情绪的1万条表达性语音数据集
- 多语言训练显著提升模型在不同语言间的泛化性能
- 模型在高低唤醒情绪间可稳定迁移,适合跨语言语音应用
语音重音在不同语言、情绪和语调风格中差异显著,但现有强调检测模型主要基于单语言中性朗读语料训练与评估。本文提出MMEE(多语言多情感强调)数据集,包含10,000条专业录制的表达性语句(总时长14.13小时),覆盖7种语言和34种情绪/风格类别,每样本配有三级感知标注(每样本10次标注)。我们在单语言、跨语言、多语言、跨情绪、跨数据集及数据规模等设置下,对两种前沿架构进行基准测试。单语言模型零样本迁移能力有限,尤其在类型学差异大的语言上性能下降明显;而多语言训练显著提升了鲁棒性。模型在高唤醒与低唤醒情绪间可稳健迁移;合成与感知基准间的双向迁移表明存在共享的重音结构;且在较小训练规模下性能仍保持稳定。
原文摘要 · Abstract (English)
Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and evaluated on monolingual neutral read speech. We introduce MMEE (Multilingual Multi-Emotion Emphasis), a corpus of 10,000 professionally recorded expressive utterances (14.13 hours) across 7 languages and 34 emotion/style categories, with three-level perceptual labels (10 annotations per sample). We benchmark two state-of-the-art architectures under monolingual, cross-lingual, multilingual, cross-emotion, cross-dataset, and data-scale settings. Monolingual models show limited zero-shot transfer, degrading across typologically distant languages, while multilingual training substantially improves robustness. Models transfer robustly between high- and low-arousal emotions; bidirectional transfer between synthetic and perceptual benchmarks suggests shared prosodic structure; and performance stays robust even at smaller training scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。