实时零样本语音转换,低延迟高质量
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
- 用发音特征空间分离语音内容与说话人特征
- 合成质量媲美当前最好方法,延迟仅61.4毫秒
- 适合需要低延迟语音转换的实时应用
语音转换在辅助沟通、娱乐等领域具有重要意义。本文提出RT-VC,一种零样本实时语音转换系统,实现超低延迟与高质量合成。该方法利用发音特征空间自然解耦语音内容与说话人特征,提升转换鲁棒性与可解释性。同时,通过可微数字信号处理(DDSP)直接从发音特征生成语音,显著降低转换延迟。实验表明,尽管合成质量与当前最佳方法(SOTA)相当,RT-VC在CPU上实现61.4毫秒延迟,较此前方法降低13.3%。
原文摘要 · Abstract (English)
Voice conversion has emerged as a pivotal technology in numerous applications ranging from assistive communication to entertainment. In this paper, we present RT-VC, a zero-shot real-time voice conversion system that delivers ultra-low latency and high-quality performance. Our approach leverages an articulatory feature space to naturally disentangle content and speaker characteristics, facilitating more robust and interpretable voice transformations. Additionally, the integration of differentiable digital signal processing (DDSP) enables efficient vocoding directly from articulatory features, significantly reducing conversion latency. Experimental evaluations demonstrate that, while maintaining synthesis quality comparable to the current state-of-the-art (SOTA) method, RT-VC achieves a CPU latency of 61.4 ms, representing a 13.3\% reduction in latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。