arXiv:2511.08496cs.SDcs.AI2025-11AAAI被引 4

提出高效高保真零样本歌声转换框架,适配资源受限场景

HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios

  • 联合提取音色与内容特征,避免信息丢失
  • 通过音高音量建模提升声音保真度,效果优于当前最优方法
  • 兼具语音超分辨率能力,适合音乐生成与低资源应用

零样本歌声转换(SVC)可在不微调的情况下,将源歌手音色转为未见过的目标歌手声音,同时保留旋律内容。现有方法分别建模音色与演唱内容,导致关键声学信息丢失且计算开销大。为此,我们提出HQ-SVC,一种高效高质的零样本歌声转换框架。HQ-SVC首先使用解耦编码器联合提取内容与说话人特征;随后通过音高和音量建模增强保真度,保留传统分离建模中易丢失的关键声学信息,并利用可微信号处理与扩散技术逐步优化输出。评估表明,HQ-SVC在转换质量与效率上显著优于当前最优零样本SVC方法。此外,其生成语音自然度超越专用音频超分辨率方法,且原生支持语音超分辨率任务。

原文摘要 · Abstract (English)

Zero-shot singing voice conversion (SVC) transforms a source singer's timbre to an unseen target speaker's voice while preserving melodic content without fine-tuning. Existing methods model speaker timbre and vocal content separately, losing essential acoustic information that degrades output quality while requiring significant computational resources. To overcome these limitations, we propose HQ-SVC, an efficient framework for high-quality zero-shot SVC. HQ-SVC first extracts jointly content and speaker features using a decoupled codec. It then enhances fidelity through pitch and volume modeling, preserving critical acoustic information typically lost in separate modeling approaches, and progressively refines outputs via differentiable signal processing and diffusion techniques. Evaluations confirm HQ-SVC significantly outperforms state-of-the-art zero-shot SVC methods in conversion quality and efficiency. Beyond voice conversion, HQ-SVC achieves superior voice naturalness compared to specialized audio super-resolution methods while natively supporting voice super-resolution tasks.

歌声转换零样本语音增强高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。