arXiv:2505.04639cs.CLcs.AI2025-05

用扩散模型同时实现语音翻译和口音转换,提升跨文化沟通自然度。

Language translation, and change of accent for speech-to-speech task using diffusion model

  • 将语音翻译与口音转换统一为条件生成任务,基于音素和目标特征生成语音。
  • 利用扩散模型生成高质量目标语音谱图,实现语言内容与口音的联合优化。
  • 相比传统流水线方法更高效,适合多语言语音交互系统开发。

语音到语音翻译(S2ST)旨在将一种语言的口语输入转换为另一种语言的口语输出,通常聚焦于语言翻译或口音适应。然而,有效的跨文化沟通需要同时处理这两方面:在翻译内容的同时,将说话人口音调整为目标语言语境。本文提出一种统一方法,实现语音翻译与口音转换的同步处理,这一任务在现有文献中仍较少被探索。我们的方法将问题重构为条件生成任务,以音素为基础,结合目标语音特征生成目标语音。利用扩散模型在高保真生成方面的优势,我们借鉴文本到图像扩散策略,通过源语音转录作为条件,生成包含期望语言与口音属性的目标语音梅尔频谱图。该集成框架实现了翻译与口音适应的联合优化,相比传统流水线方法具有更高的参数效率和更强的表现力。

原文摘要 · Abstract (English)

Speech-to-speech translation (S2ST) aims to convert spoken input in one language to spoken output in another, typically focusing on either language translation or accent adaptation. However, effective cross-cultural communication requires handling both aspects simultaneously - translating content while adapting the speaker's accent to match the target language context. In this work, we propose a unified approach for simultaneous speech translation and change of accent, a task that remains underexplored in current literature. Our method reformulates the problem as a conditional generation task, where target speech is generated based on phonemes and guided by target speech features. Leveraging the power of diffusion models, known for high-fidelity generative capabilities, we adapt text-to-image diffusion strategies by conditioning on source speech transcriptions and generating Mel spectrograms representing the target speech with desired linguistic and accentual attributes. This integrated framework enables joint optimization of translation and accent adaptation, offering a more parameter-efficient and effective model compared to traditional pipelines.

语音翻译扩散模型口音转换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。