构建21小时卢森堡语情感语音数据集,填补低资源语言研究空白
LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish

- 半自动流程整合语音检测、去噪、情绪预测与人工校验
- 涵盖4类情感,支持卢森堡语表达性语音合成基准测试
- 适合低资源语言语音技术研究者参考
当前最先进的语音数据集多聚焦于主流语言,常忽视卢森堡语等低资源语言。本文提出LuxEmo,一个21小时的卢森堡语对话式情感语音语料库,包含4种情感类别。该语料库源自Radio Télévision Luxembourg(RTL)青年广播节目,采用自动化检测结合人工验证的方式构建。我们提出一种半自动清洗流程,融合语音活动检测、去噪、语言识别、LuxASR分段、自动情绪预测、词汇线索及定向人工审查。此外,我们对五种表达性语音合成系统进行了基准测试,涵盖基于德语的跨语言迁移、多语言卢森堡语支持、卢森堡语适配和非参数化韵律迁移。性能通过客观指标和人工评估进行衡量。
原文摘要 · Abstract (English)
State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low-resource languages such as Luxembourgish, which remain underrepresented in speech technology research. In this work, we introduce LuxEmo, a 21-hour conversational expressive speech corpus for Luxembourgish with 4 emotion categories. LuxEmo is derived from Radio Télévision Luxembourg (RTL) youth broadcasts, using automated detection followed by human validation. We propose a semi-automatic curation workflow combining voice activity detection, denoising, language identification, LuxASR-based segmentation, automatic emotion prediction, lexical cues, and targeted human review. Additionally, we benchmark five expressive TTS systems covering German-based cross-lingual transfer, multilingual Luxembourgish support, Luxembourgish adaptation, and non-parametric prosody transfer. Performance is evaluated using both objective metrics and human evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。