arXiv:2409.08583cs.SDcs.AI2024-09中稿 · ICASSP 2025被引 6

轻量级歌声转换模型,高效低耗且保持高音质。

LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling

  • 基于扩散模型构建轻量化框架,适配CPU运行
  • 推理速度提升显著,多设备表现优异
  • 适合资源受限场景下的高质量歌声转换

歌声转换(SVC)作为语音转换的子领域,可将一名歌手的声音转换为另一名歌手的声音,同时保留旋律、节奏和音色等音乐元素。传统SVC方法在音频质量、数据需求和计算复杂度方面存在局限。本文提出LHQ-SVC,一种基于SVC框架与扩散模型的轻量级、兼容CPU的模型,旨在降低模型尺寸与计算开销,同时保持性能。通过引入优化特征提升推理质量,并利用性能调优工具与并行计算框架优化CPU执行效率。实验表明,LHQ-SVC在不同设备上均实现显著的处理速度与效率提升,具有竞争力的转换性能,适用于资源受限环境下的高质量歌声转换应用。

原文摘要 · Abstract (English)

Singing Voice Conversion (SVC) has emerged as a significant subfield of Voice Conversion (VC), enabling the transformation of one singer's voice into another while preserving musical elements such as melody, rhythm, and timbre. Traditional SVC methods have limitations in terms of audio quality, data requirements, and computational complexity. In this paper, we propose LHQ-SVC, a lightweight, CPU-compatible model based on the SVC framework and diffusion model, designed to reduce model size and computational demand without sacrificing performance. We incorporate features to improve inference quality, and optimize for CPU execution by using performance tuning tools and parallel computing frameworks. Our experiments demonstrate that LHQ-SVC maintains competitive performance, with significant improvements in processing speed and efficiency across different devices. The results suggest that LHQ-SVC can meet

歌声转换轻量模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。