arXiv:2505.18609cs.CL2025-05中稿 · Interspeech 2025被引 9

打造23种印度语语音数据集,实现精准可控的语音合成。

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations

  • 构建含1.3万小时语音与2400万标注的多属性数据集。
  • 首个开源支持文本描述引导的印度语语音合成模型。
  • 可跨语言迁移情感与语调特征,适合多语种语音开发。

我们提出RASMALAI,一个大规模语音数据集,包含13,000小时语音和2400万条细粒度文本描述标注,涵盖说话人身份、口音、情绪、风格及背景条件等属性,旨在推动23种印度语言及英语的可控、富有表现力的文本到语音(TTS)合成。基于此,我们开发了IndicParlerTTS,首个开源的文本描述引导式印度语TTS模型。系统评估表明,该模型能高质量生成特定说话人语音,准确遵循文本描述并精确合成指定属性。此外,其在语言内与跨语言间均有效迁移表达特征。IndicParlerTTS在各项评测中表现优异,为印度语可控多语言语音合成树立新标准。

原文摘要 · Abstract (English)

We introduce RASMALAI, a large-scale speech dataset with rich text descriptions, designed to advance controllable and expressive text-to-speech (TTS) synthesis for 23 Indian languages and English. It comprises 13,000 hours of speech and 24 million text-description annotations with fine-grained attributes like speaker identity, accent, emotion, style, and background conditions. Using RASMALAI, we develop IndicParlerTTS, the first open-source, text-description-guided TTS for Indian languages. Systematic evaluation demonstrates its ability to generate high-quality speech for named speakers, reliably follow text descriptions and accurately synthesize specified attributes. Additionally, it effectively transfers expressive characteristics both within and across languages. IndicParlerTTS consistently achieves strong performance across these evaluations, setting a new standard for controllable multilingual expressive speech synthesis in Indian languages.

语音合成多语言数据集可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。