arXiv:2605.12310cs.SD2026-05中稿 · ICASSP 2026

解决伴奏中和声干扰的歌声转换,让合成歌声更自然。

Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling

论文配图:Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling
图 1 · 摘自论文原文
  • 用CQT提取器同时保留主旋律和残留和声,避免信息丢失。
  • 在含丰富和声与单旋律录音上,音色相似度与自然度均优于基线模型。
  • 零样本跨语言适用,适合需要高质量多声部歌声生成的研究者。

歌声转换(SVC)旨在将源演唱声音转换为目标歌手的声音,同时保持歌词和旋律。现有方法依赖于音高提取器从纯净人声中捕捉主旋律,但无法可靠地从伴奏录音中提取纯净人声,常留下残余和声。本文提出Poly-SVC,一种零样本、跨语言的歌声转换系统,可处理残余和声。该系统包含三个核心组件:基于常数Q变换(CQT)的音高提取器,用于同时保留主旋律与残余和声;随机采样器,以减少CQT中的干扰信息;以及基于条件流匹配(CFM)的扩散解码器,融合音高、内容与音色特征,生成自然的多声部输出。实验表明,Poly-SVC在含丰富和声与单旋律录音上,均显著优于基线模型,在自然度、音色相似度与和声重建方面表现更佳。

原文摘要 · Abstract (English)

Singing Voice Conversion (SVC) aims to transform a source singing voice into a target singer while preserving lyrics and melody. Most existing SVC methods depend on F0 extractors to capture the lead melody from clean vocals. However, no existing method can reliably extract clean vocals from accompanied recordings without leaving residual harmonies behind. In this paper, we innovatively propose Poly-SVC, a zero-shot, cross-lingual singing voice conversion system designed to process residual harmonies. Poly-SVC is composed of three key components: a Constant-Q Transform (CQT)-based pitch extractor to preserve both the lead melody and residual harmony, a random sampler to reduce interference information from the CQT and a diffusion decoder based on Conditional Flow Matching (CFM) that fuses pitch, content, and timbre features into natural-sounding polyphonic outputs. Experiments demonstrate that Poly-SVC surpasses the baseline models in naturalness, timbre similarity and harmony reconstruction across both harmony-rich and single-melody recordings.

歌声转换多声部扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。