用合成数据训练语音口音归一化模型,实现自然度与时长控制的平衡。
CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data
- 用合成源语音+真实母语语音构建训练数据,避免TTS伪影
- 无需真实非母语语音即可训练,仍保持高内容保真度
- 非自回归模型支持灵活节奏控制,适合口音迁移应用
口音归一化系统常因训练数据不佳和时长建模僵化导致输出不自然且内容失真。本文提出一种‘源语音合成’的数据构建方法:通过生成源语言(L2)语音,并以真实母语语音作为训练目标,避免学习到语音合成(TTS)的伪影,且训练过程中无需真实非母语数据。同时,我们提出CosyAccent——一种非自回归模型,解决了韵律自然性与时长控制之间的权衡问题。该模型隐式建模节奏以增强灵活性,同时提供对总输出时长的显式控制。实验表明,尽管未使用任何真实非母语语音进行训练,CosyAccent在内容保真度上显著优于强基线模型,并在自然度方面表现更优。
原文摘要 · Abstract (English)
Accent normalization (AN) systems often struggle with unnatural outputs and undesired content distortion, stemming from both suboptimal training data and rigid duration modeling. In this paper, we propose a "source-synthesis" methodology for training data construction. By generating source L2 speech and using authentic native speech as the training target, our approach avoids learning from TTS artifacts and, crucially, requires no real L2 data in training. Alongside this data strategy, we introduce CosyAccent, a non-autoregressive model that resolves the trade-off between prosodic naturalness and duration control. CosyAccent implicitly models rhythm for flexibility yet offers explicit control over total output duration. Experiments show that, despite being trained without any real L2 speech, CosyAccent achieves significantly improved content preservation and superior naturalness compared to strong baselines trained on real-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。