用文本转写生成多口音语音,提升语音风格转换效果
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
- 通过大模型生成转写文本,驱动多语言语音合成生成口音样本
- 在母语与非母语者上验证,显著提升口音转换的自然度和保真度
- 适合语音合成、跨口音研究及个性化语音系统开发者
在带口音的语音转换中,目标是将一种口音的语音转换为另一种,同时保持说话人身份和语义内容。本文提出一种新方法,通过文本转写生成同一说话人的多口音语音对,用于训练口音转换系统。首先利用大语言模型(LLMs)生成转写文本,再输入多语言文本转语音(TTS)模型合成带有不同口音的英语语音。作为对比基线,我们在合成的平行语料上构建了一个序列到序列模型进行口音转换。实验验证了该方法在母语和非母语英语说话者上的有效性。主观与客观评估进一步证明该数据集在口音转换研究中的价值。
原文摘要 · Abstract (English)
In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same speaker, through text transliteration for training accent conversion systems. We begin by generating transliterated text with Large Language Models (LLMs), which is then fed into multilingual TTS models to synthesize accented English speech. As a reference system, we built a sequence-to-sequence model on the synthetic parallel corpus for accent conversion. We validated the proposed method for both native and non-native English speakers. Subjective and objective evaluations further validate our dataset's effectiveness in accent conversion studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。