用合成数据训练端到端语音转换,实现低延迟实时通话
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
- 用预训练零样本模型生成合成数据,直接学习音色转换
- 端到端延迟仅77.1毫秒,自然度和音色相似性更优
- 无需语音识别或声纹分离模块,适合实时语音应用
语音转换(VC)旨在保留语言内容的同时改变说话人音色。尽管近期VC模型性能优异,但多数在实时流式场景中表现不佳,受限于高延迟、依赖语音识别模块或复杂的声纹解耦,常导致音色泄漏或自然度下降。我们提出SynthVC,一种基于神经音频编解码架构的端到端流式语音转换框架,通过预训练的零样本语音转换模型生成的合成平行数据直接学习音色转换,无需显式的语义-声纹分离或识别模块。实验表明,SynthVC在自然度和音色相似性上优于基线流式系统,端到端延迟仅为77.1毫秒。
原文摘要 · Abstract (English)
Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming scenarios due to high latency, dependence on ASR modules, or complex speaker disentanglement, which often results in timbre leakage or degraded naturalness. We present SynthVC, a streaming end-to-end VC framework that directly learns speaker timbre transformation from synthetic parallel data generated by a pre-trained zero-shot VC model. This design eliminates the need for explicit content-speaker separation or recognition modules. Built upon a neural audio codec architecture, SynthVC supports low-latency streaming inference with high output fidelity. Experimental results show that SynthVC outperforms baseline streaming VC systems in both naturalness and speaker similarity, achieving an end-to-end latency of just 77.1 ms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。