Conan实现零样本实时语音转换,保持语义同时自然还原音色。
Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
- 分块流式处理+自适应风格编码,实时转换更稳定
- 主观评测优于基线模型,语音自然度显著提升
- 适合实时语音通信与跨说话人音色迁移场景
零样本在线语音转换在实时通信和娱乐中具有重要应用前景。然而,现有模型在实时约束下难以保持语义一致性、生成自然语音,且对未见说话人特征适应能力弱。为此,我们提出Conan:一种分块流式零样本语音转换模型,能在保留源语音内容的同时匹配参考语音的音色与风格。Conan包含三个核心组件:1)基于Emformer的流式内容提取器,实现低延迟内容编码;2)自适应风格编码器,从参考语音中提取细粒度风格特征以增强风格适配;3)全因果像素混洗声码器(Causal Shuffle Vocoder),基于像素混洗机制实现完全因果的HiFiGAN。实验表明,Conan在主观与客观指标上均优于基线模型。音频样例可访问 https://aaronz345.github.io/ConanDemo。
原文摘要 · Abstract (English)
Zero-shot online voice conversion (VC) holds significant promise for real-time communications and entertainment. However, current VC models struggle to preserve semantic fidelity under real-time constraints, deliver natural-sounding conversions, and adapt effectively to unseen speaker characteristics. To address these challenges, we introduce Conan, a chunkwise online zero-shot voice conversion model that preserves the content of the source while matching the voice timbre and styles of reference speech. Conan comprises three core components: 1) a Stream Content Extractor that leverages Emformer for low-latency streaming content encoding; 2) an Adaptive Style Encoder that extracts fine-grained stylistic features from reference speech for enhanced style adaptation; 3) a Causal Shuffle Vocoder that implements a fully causal HiFiGAN using a pixel-shuffle mechanism. Experimental evaluations demonstrate that Conan outperforms baseline models in subjective and objective metrics. Audio samples can be found at https://aaronz345.github.io/ConanDemo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。