提升语音转换实时性与鲁棒性,支持低延迟零样本语音转换
MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion

- 采用未来感知分块机制,优化时序建模结构
- 40毫秒分块下仍保持稳定转换,延迟降至110毫秒
- 新声线编码器提升对参考音频质量的鲁棒性,适合实际应用
流式零样本语音转换因在实时应用中的潜力而日益流行。最近提出的MeanVC实现了轻量级流式零样本语音转换,但存在若干局限:其分块自回归去噪使有效训练序列长度翻倍,小分块设置下转换质量下降,且声线编码器直接依赖参考梅尔谱图,对参考音频质量敏感。为此我们提出MeanVC 2。引入未来感知分块(FRC),显式调度扩散变换器解码器层中的过去与未来感受野,并移除干净分块教师强制。通过引入有界未来上下文,FRC实现40毫秒分块下的稳定转换。进一步提出通用声线标记编码器,从全局说话人嵌入构建声线表示,并通过交叉注意力检索细粒度声线线索,增强对低质量参考音频的鲁棒性及零样本说话人相似性。实验表明,MeanVC 2显著优于MeanVC,同时将延迟从211毫秒降低至110毫秒。音频样本已公开,源代码将发布。
原文摘要 · Abstract (English)
Streaming zero-shot voice conversion (VC) has become increasingly popular due to its potential for real-time applications. The recently proposed MeanVC achieves lightweight streaming zero-shot VC, but it has several limitations: its chunk-wise autoregressive denoising doubles the effective training sequence length, conversion quality degrades under small-chunk settings, and its timbre encoder directly relies on reference mel-spectrograms, making it sensitive to reference audio quality. To address these limitations we propose MeanVC 2. We introduce future-receptive chunking (FRC), which explicitly schedules past and future receptive fields across diffusion transformer decoder layers and removes clean-chunk teacher forcing. By incorporating bounded future context, FRC enables stable conversion with a 40 ms chunk size. We further introduce a universal timbre token encoder, which constructs a timbre representation from a global speaker embedding and retrieves fine-grained timbre cues via cross-attention, improving robustness to low-quality references and enhancing zero-shot speaker similarity. Experimental results show that MeanVC 2 significantly outperforms MeanVC, while reducing latency from 211 ms to 110 ms. Audio samples are publicly available. The source code will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。