通过边界感知瓶颈提升歌声风格转换的精细控制与保真度
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck
- 用边界感知的语音瓶颈提取音素区间特征,抑制残留源风格
- 引入帧级技术矩阵和目标音高处理,实现稳定动态风格渲染
- 利用48kHz辅助模型补全高频频谱,缓解数据稀缺问题
本文提出S4团队参加2025年歌声转换挑战赛(SVCC2025)的新型歌声风格转换系统,专注于域内设置下的细粒度风格控制。为应对风格泄漏、动态渲染不稳及小样本下高保真生成的挑战,提出三项创新:边界感知的Whisper瓶颈,通过聚合音素跨度表示,抑制残留源风格并保留语言内容;显式的帧级技术矩阵,结合推理时的定向音高处理,实现稳定且清晰的动态风格渲染;以及基于听觉感知的高频带补全策略,利用辅助标准48kHz SVC模型增强高频谱,克服数据稀缺问题而不产生过拟合。在官方主观评估中,本系统在自然度上表现最佳,同时在说话人相似性和技术控制性上保持竞争力,且使用的额外歌唱数据远少于其他顶尖系统。音频样例可在线获取。
原文摘要 · Abstract (English)
This paper presents the submission of the S4 team to the Singing Voice Conversion Challenge 2025 (SVCC2025)-a novel singing style conversion system that advances fine-grained style conversion and control within in-domain settings. To address the critical challenges of style leakage, dynamic rendering, and high-fidelity generation with limited data, we introduce three key innovations: a boundary-aware Whisper bottleneck that pools phoneme-span representations to suppress residual source style while preserving linguistic content; an explicit frame-level technique matrix, enhanced by targeted F0 processing during inference, for stable and distinct dynamic style rendering; and a perceptually motivated high-frequency band completion strategy that leverages an auxiliary standard 48kHz SVC model to augment the high-frequency spectrum, thereby overcoming data scarcity without overfitting. In the official SVCC2025 subjective evaluation, our system achieves the best naturalness performance among all submissions while maintaining competitive results in speaker similarity and technique control, despite using significantly less extra singing data than other top-performing systems. Audio samples are available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。