提升语音翻译中说话人特征保留,实现高保真非自回归翻译。
Preserving Speaker Information in Direct Speech-to-Speech Translation with Non-Autoregressive Generation and Pretraining
- 引入自监督预训练与特征融合策略,增强说话人信息提取。
- 在CVSS-T数据集上BLEU提升1.14,说话人相似性显著提高。
- 保持低延迟翻译(每句仅增0.04秒),适合实时多语言沟通场景。
语音到语音翻译(S2ST)旨在将一种语言的语音转换为语义等价的另一种语言语音,促进不同语言使用者间的交流。主流的语音到离散单元翻译(S2UT)方法可缓解传统级联系统中的误差传播和推理慢问题。然而,由于离散单元主要捕获内容信息,现有S2UT方法难以保留源语音的说话人特征。我们此前提出的SC-S2UT通过引入说话人适配器和单元到梅尔谱图结构,实现了说话人信息保留与非自回归生成。本研究在此基础上提出一种自监督预训练方法,以丰富说话人适配器和单元到梅尔结构所提取的信息,并探索多种特征融合策略优化说话人与内容特征的整合。在CVSS-T数据集上的ES-EN和FR-EN任务实验表明,该方法相较SC-S2UT BLEU提升1.14,同时显著改善了平均意见分(MOS)和说话人相似性。此外,翻译质量接近传统S2UT,每句推理时间仅增加0.04秒,且保持高说话人相似性。结果验证了方法的有效性。
原文摘要 · Abstract (English)
Speech-to-Speech Translation (S2ST) refers to the conversion of speech in one language into semantically equivalent speech in another language, facilitating communication between speakers of different languages. Speech-to-Discrete Unit Translation (S2UT), a mainstream approach for end-to-end S2ST, addresses challenges such as error propagation across modules and slow inference speed often encountered in traditional cascade systems. However, as discrete units primarily capture content information, conventional S2UT methods fail to retain speaker-specific characteristics from the source. Our previous work, SC-S2UT, introduced a speaker adapter and a unit-to-mel structure, enabling the preservation of speaker information and non-autoregressive speech generation. Building on this foundation, this study proposes a self-supervised pretraining method to enrich the information extracted by both the speaker adapter and the unit-to-mel structure. Additionally, we investigate different feature fusion strategies to further improve the integration of speaker and content features. Experiments conducted on the CVSS-T dataset for ES-EN and FR-EN tasks demonstrate that our proposed method achieves a BLEU score improvement of 1.14 compared to SC-S2UT, along with significant enhancements in MOS and speaker similarity. Furthermore, our approach achieves translation quality comparable to traditional S2UT, with only a minimal increase of 0.04s per utterance in inference time, while maintaining high speaker similarity. These results validate the effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。