用自适应模块分离语音内容与风格,实现高质量零样本语音转换
AdaptVC: High Quality Voice Conversion with Adaptive Learning
- 通过适配器动态提取自监督特征中的语音内容与风格
- 零样本测试中语音质量与参考音色相似度均超越现有方法
- 适合需要高保真语音转换的语音合成与个性化应用
语音转换的目标是将源说话人的语音转换为参考说话人音色,同时保留原始语义内容。核心挑战在于如何有效分离语言内容与说话人风格。现有方法虽尝试分离二者,但泛化能力仍需提升,尤其在零样本场景下表现不足。本文通过适配器微调自监督语音特征,成功实现内容与说话人特征的解耦。适配器动态编码来自丰富自监督特征的细微差异,解码器融合这些特征以生成高度匹配参考音色且内容损失最小的语音。此外,采用带交叉注意力的条件流匹配解码器进一步提升合成质量与效率。在零样本场景下的主观与客观评估表明,该方法在语音质量与参考语音相似度上优于现有模型。
原文摘要 · Abstract (English)
The goal of voice conversion is to transform the speech of a source speaker to sound like that of a reference speaker while preserving the original content. A key challenge is to extract disentangled linguistic content from the source and voice style from the reference. While existing approaches leverage various methods to isolate the two, a generalization still requires further attention, especially for robustness in zero-shot scenarios. In this paper, we achieve successful disentanglement of content and speaker features by tuning self-supervised speech features with adapters. The adapters are trained to dynamically encode nuanced features from rich self-supervised features, and the decoder fuses them to produce speech that accurately resembles the reference with minimal loss of content. Moreover, we leverage a conditional flow matching decoder with cross-attention speaker conditioning to further boost the synthesis quality and efficiency. Subjective and objective evaluations in a zero-shot scenario demonstrate that the proposed method outperforms existing models in speech quality and similarity to the reference speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。