arXiv:2501.02953cs.SDeess.AS2025-01中稿 · ICASSP 2025被引 5

提升歌声转换质量,引入后处理和专业评测集

SYKI-SVC: Advancing Singing Voice Conversion with Post-Processing Innovations and an Open-Source Professional Testset

  • 融合内容向量与语音模型提取音高与语义特征
  • 后处理增强高频信息,显著改善音质清晰度
  • 首个面向专业表达的公开评测数据集,适合音乐生成研究者

歌声转换旨在将源演唱声线转换为目标歌手声线,同时保留原歌词、旋律及多种演唱技巧。本文提出一种高保真歌声转换系统,基于SVCC T02框架,包含特征提取器、声线转换器和后处理器三部分。特征提取器利用ContentVec和Whisper模型提取输入歌声的基频轮廓与非说话人依赖的语义特征。声线转换器整合提取的音色、基频与语言内容,合成目标歌手波形。后处理器通过简单有效的信号处理,直接从源信号中增强高频信息,提升音频质量。由于缺乏标准化的专业评测数据集,本文构建并公开了一个专用测试集。对比评估表明,本系统在自然度方面表现优异,分析进一步验证了系统设计的有效性。

原文摘要 · Abstract (English)

Singing voice conversion aims to transform a source singing voice into that of a target singer while preserving the original lyrics, melody, and various vocal techniques. In this paper, we propose a high-fidelity singing voice conversion system. Our system builds upon the SVCC T02 framework and consists of three key components: a feature extractor, a voice converter, and a post-processor. The feature extractor utilizes the ContentVec and Whisper models to derive F0 contours and extract speaker-independent linguistic features from the input singing voice. The voice converter then integrates the extracted timbre, F0, and linguistic content to synthesize the target speaker's waveform. The post-processor augments high-frequency information directly from the source through simple and effective signal processing to enhance audio quality. Due to the lack of a standardized professional dataset for evaluating expressive singing conversion systems, we have created and made publicly available a specialized test set. Comparative evaluations demonstrate that our system achieves a remarkably high level of naturalness, and further analysis confirms the efficacy of our proposed system design.

歌声转换后处理音质增强数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。