仅用单语语料提升大模型的混语语音合成能力
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
- 通过拼接不同单语语料中的词构建混语数据
- 在有限数据下实现更自然、一致的混语语音合成
- 适合需要多语言混合场景的语音系统开发者
尽管大型语言模型在语音生成与识别方面展现出潜力,但其应用仍主要局限于单语场景,对混语(CS)情境的探索有限。本文提出一种混语大语言模型(CS-LLM),仅使用单语语料即可增强大模型的混语文本转语音(CS TTS)能力。具体而言,首先通过多语言语音识别与合成任务提升模型的多语处理能力;随后设计了一种有效的混语数据构建策略,将不同单语语音语料中的词汇拆分并拼接,使模型获得更强的混语语音合成能力。实验表明,该方法在自然度、说话人一致性与相似性方面均优于基线,即使数据有限也表现优异。此外,所构建的混语数据还进一步提升了多语言语音合成与识别性能。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have shown potential in speech generation and recognition, their applications are mainly confined to monolingual scenarios, with limited explorations in code-switched (CS) contexts. In this paper, we propose a Code-Switched Large Language Model (CS-LLM) to enhance the code-switched text-to-speech synthesis (CS TTS) capability in LLMs with only monolingual corpora. Specifically, we begin by enhancing the multilingual speech processing ability of LLMs through multilingual speech recognition and synthesis tasks. Then, we develop an effective code-switched (CS) data construction strategy that splits and concatenates words from different monolingual speech corpora to equip LLMs with improved CS TTS ability. Experiments show that our approach outperforms baselines in CS TTS in terms of naturalness, speaker consistency and similarity even with limited data. Additionally, the constructed CS data further improves multilingual speech synthesis and recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。