arXiv:2504.05686cs.SDcs.AI2025-04被引 5

提升零样本歌声转换鲁棒性,解决音色平淡与拼接不顺问题

kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization

  • 用加法合成融合音高轮廓与频谱,改善WavLM表征的谐波缺失
  • 提出新距离度量与加权优化策略,显著提升音轨拼接流畅度
  • 方法通用性强,适用于各类基于拼接的神经声学模型

零样本歌声转换(SVC)的鲁棒性至关重要。本文针对kNN-VC框架提出两项新方法以增强其鲁棒性。首先,kNN-VC的核心表征WavLM缺乏谐波强调,导致音色沉闷并产生嗡鸣伪影。为此,我们利用WavLM、音高轨迹与频谱间的双射关系,实施加法合成,并将生成波形融入模型以缓解该问题。其次,kNN-VC忽略拼接平滑性这一关键感知因素。为此,我们设计了一种新距离度量,筛选不适配的kNN候选,并在推理阶段优化候选加权和。尽管技术实现基于kNN-VC框架,但其适用范围广泛,可推广至一般拼接式神经合成模型。实验验证了所提方法在实现鲁棒SVC方面的有效性。演示:http://knnsvc.com 代码:https://github.com/SmoothKen/knn-svc

原文摘要 · Abstract (English)

Robustness is critical in zero-shot singing voice conversion (SVC). This paper introduces two novel methods to strengthen the robustness of the kNN-VC framework for SVC. First, kNN-VC's core representation, WavLM, lacks harmonic emphasis, resulting in dull sounds and ringing artifacts. To address this, we leverage the bijection between WavLM, pitch contours, and spectrograms to perform additive synthesis, integrating the resulting waveform into the model to mitigate these issues. Second, kNN-VC overlooks concatenative smoothness, a key perceptual factor in SVC. To enhance smoothness, we propose a new distance metric that filters out unsuitable kNN candidates and optimize the summing weights of the candidates during inference. Although our techniques are built on the kNN-VC framework for implementation convenience, they are broadly applicable to general concatenative neural synthesis models. Experimental results validate the effectiveness of these modifications in achieving robust SVC. Demo: http://knnsvc.com Code: https://github.com/SmoothKen/knn-svc

歌声转换音频合成拼接优化零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。