无需音标转换,直接从多文字混合文本生成自然语音。
MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
- 用预训练语音自监督模型+T5编码器,从混合文字直接生成伪语言标签。
- 在无音标转换情况下,语音合成效果媲美传统方法,保留语调和口音特征。
- 适合大规模非标注语音数据,降低人工转写成本,提升可扩展性。
本研究提出一种新型语音合成方法,可替代传统的字符到音素(G2P)转换。通过深度学习模型,直接从语音中生成离散标记。利用预训练的语音自监督学习(SSL)模型,训练T5编码器从包含汉字与假名的混合脚本文本中生成伪语言标签。该方法免去人工音素标注,显著降低制作成本并提升可扩展性,尤其适用于大规模未标注语音数据集。模型性能与传统基于G2P的文本到语音系统相当,能合成保留自然语言及韵律特征(如口音、语调)的语音。
原文摘要 · Abstract (English)
This study presents a novel approach to voice synthesis that can substitute the traditional grapheme-to-phoneme (G2P) conversion by using a deep learning-based model that generates discrete tokens directly from speech. Utilizing a pre-trained voice SSL model, we train a T5 encoder to produce pseudo-language labels from mixed-script texts (e.g., containing Kanji and Kana). This method eliminates the need for manual phonetic transcription, reducing costs and enhancing scalability, especially for large non-transcribed audio datasets. Our model matches the performance of conventional G2P-based text-to-speech systems and is capable of synthesizing speech that retains natural linguistic and paralinguistic features, such as accents and intonations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。