跨语言语音克隆让说话人用目标语言发声,保持原音色并提升可懂度。
KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026
- 用语言标签提示增强多语言控制,减少口音泄露。
- 强化学习微调使语音可懂度提升,尤其在复杂语境下。
- 参考语音匹配专有名词发音,适合术语密集场景使用。
跨语言语音克隆旨在以目标语言生成语音,同时保留源语言参考语音的说话人身份,是语音翻译的核心任务,也是 IWSLT 2026 跨语言语音克隆赛道的重点。主要挑战在于应对口音差异和领域专有词汇带来的可懂度与自然度下降问题。我们基于多语言文语转换模型 FishAudio-S2-Pro,引入语言标签提示以增强语言控制、减少口音泄露;进一步采用强化学习(RL)微调进行任务适配,观察到可懂度显著提升;最后提出一种参考条件下的词汇匹配方法,在存在词汇重叠时改善领域专有术语的发音。实验结果表明,语言提示带来最大性能增益,而词汇匹配在匹配子集上表现出持续改进效果。
原文摘要 · Abstract (English)
Cross-lingual voice cloning aims to generate speech in a target language while preserving speaker identity from a source-language reference. This task is central to speech translation and is the focus of the IWSLT 2026 Cross-Lingual Voice Cloning track. A key challenge is maintaining intelligibility and naturalness in the presence of accent variation and domain-specific vocabulary. We build on a multilingual text-to-speech model, FishAudio-S2-Pro, and introduce language tag prompting to improve language control and reduce accent leakage. We further apply reinforcement learning (RL) fine-tuning for task adaptation and observe improvements in intelligibility. Finally, we propose a reference-conditioned lexical matching method that improves pronunciation of domain-specific terms when lexical overlap is present. Results show that language prompting provides the largest gains, while lexical matching yields consistent improvements on matched subsets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。