arXiv:2409.18622cs.SDeess.AS2024-09EMNLP被引 2

从音频直接提取语言特征,提升多语言低资源语音合成效果

Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech

论文配图:Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
图 1 · 摘自论文原文
  • 通过音频直接提取语言特征,过滤说话人音色等无关信息
  • 在未见过的语言上实现有效迁移,低资源场景表现更优
  • 适合多语言语音合成与数据稀缺场景的应用研究

高质量数据的获取困难,尤其在多语言环境下,促使研究者关注低资源场景。现有方法依赖固定的语言标识符表达,导致语言表征学习不足,难以生成未见语言的语音。为此,我们提出一种新方法:直接从音频输入中提取语言特征,同时有效过滤掉包括音色在内的非语言声学信息。主观与客观评估验证了该方法在多语言语音合成中的有效性,并凸显其在未见语言上的低资源迁移学习优势。

原文摘要 · Abstract (English)

The difficulty of acquiring abundant, high-quality data, especially in multi-lingual contexts, has sparked interest in addressing low-resource scenarios. Moreover, current literature rely on fixed expressions from language IDs, which results in the inadequate learning of language representations, and the failure to generate speech in unseen languages. To address these challenges, we propose a novel method that directly extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. Subjective and objective evaluations affirm the effectiveness of our approach for multi-lingual text-to-speech, and highlight its superiority in low-resource transfer learning for previously unseen language.

语音合成多语言低资源音频特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。