一个模型搞定语音转换三大场景,零样本统一建模。
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
- 用MoE结构分离通用知识与场景特异性表达。
- 三类场景表现均超越专用模型,2步即可快速生成。
- 适合需要多场景统一语音转换的研究与应用。
语音转换领域在说话人克隆和语言保持方面取得进展,但当前仍依赖针对语言保持、情感表达和歌唱场景的专用模型。本文提出OneVoice,一种统一的零样本框架,可在单一模型中处理全部三种场景。该框架基于无变分自编码器的下一补丁扩散连续语言模型,确保高保真度和高效序列建模。其核心设计采用混合专家(MoE)结构,显式建模共享转换知识与场景特异性表达力。专家选择由双路径路由机制协调,包括共享专家隔离与基于全局-局部线索的场景感知域专家分配。为实现精准控制,各层融合场景特定的韵律特征,通过门控机制实现韵律信息的自适应使用。为支持核心思想并缓解数据不平衡(语音数据丰富而歌唱数据稀缺),采用两阶段渐进训练:基础预训练与基于LoRA的域专家场景增强。实验表明,OneVoice在所有三个场景中的表现均达到或超过专用模型水平,同时支持灵活场景切换,并提供仅需2步的快速解码版本。音频样例可在演示页面获取。
原文摘要 · Abstract (English)
Recent progress of voice conversion~(VC) has achieved a new milestone in speaker cloning and linguistic preservation. But the field remains fragmented, relying on specialized models for linguistic-preserving, expressive, and singing scenarios. We propose OneVoice, a unified zero-shot framework capable of handling all three scenarios within a single model. OneVoice is built upon a continuous language model trained with VAE-free next-patch diffusion, ensuring high fidelity and efficient sequence modeling. Its core design for unification lies in a Mixture-of-Experts (MoE) designed to explicitly model shared conversion knowledge and scenario-specific expressivity. Expert selection is coordinated by a dual-path routing mechanism, including shared expert isolation and scenario-aware domain expert assignment with global-local cues. For precise conditioning, scenario-specific prosodic features are fused into each layer via a gated mechanism, allowing adaptive usage of prosody information. Furthermore, to enable the core idea and alleviate the imbalanced issue (abundant speech vs. scarce singing), we adopt a two-stage progressive training that includes foundational pre-training and scenario enhancement with LoRA-based domain experts. Experiments show that OneVoice matches or surpasses specialized models across all three scenarios, while verifying flexible control over scenarios and offering a fast decoding version as few as 2 steps. Audio samples are available on demo page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。