利用量化残差实现零样本语音转换中的说话人与语义解耦
QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion
- 通过分析语音的时序特性,挖掘量化残差中的发音和语调信息
- 在无监督条件下实现高保真语音转换,提升可懂度与语调保留
- 适合追求高质量语音转换且不想复杂训练的开发者使用
零样本语音转换仅需单个参考语音即可将输入语音的说话人身份转换为目标说话人,无需额外训练。现有方法广泛采用自监督学习特征结合K-means量化提取高质量内容表示并去除说话人信息,但该过程也消除了细粒度的发音和语调变化,导致可懂度和语调保留下降。以往工作主要关注量化表示,而量化残差仍未被充分挖掘。本文提出一种新方法,充分利用量化残差的时序特性,实现说话人身份与语音内容的线性解耦,并恢复量化中丢失的发音和语调细节。仅使用K-means量化和线性投影,无需复杂结构或显式监督,即可实现简单高效的解耦,仅依赖重建损失进行训练。实验表明,该模型在主观与客观指标上均优于现有方法,显著提升可懂度、说话人相似度及语调保持能力,验证了线性解耦模块的有效性。
原文摘要 · Abstract (English)
Zero-shot voice conversion is a technique that alters the speaker identity of an input speech to match a target speaker using only a single reference utterance, without requiring additional training. Recent approaches extensively utilize self-supervised learning features with K-means quantization to extract high-quality content representations while removing speaker identity. However, this quantization process also eliminates fine-grained phonetic and prosodic variations, degrading intelligibility and prosody preservation. While prior works have primarily focused on quantized representations, quantization residuals remain underutilized and deserve further exploration. In this paper, we introduce a novel approach that fully utilizes quantization residuals by leveraging temporal properties of speech components. This facilitates the disentanglement of speaker identity and the recovery of phonetic and prosodic details lost during quantization. By applying only K-means quantization and linear projections, our method achieves simple yet effective disentanglement, without requiring complex architectures or explicit supervision. This allows for high-fidelity voice conversion trained solely with reconstruction losses. Experiments show that the proposed model outperforms existing methods across both subjective and objective metrics. It achieves superior intelligibility and speaker similarity, along with improved prosody preservation, highlighting the impact of our Linear Disentangler module.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。