arXiv:2601.13629eess.ASeess.SP2026-01中稿 · ICASSP 2026被引 3

S²Voice实现高保真歌声风格转换,支持跨风格迁移与零样本适配。

S$^2$Voice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion

  • 通过风格嵌入与交叉注意力增强模型对演唱风格的精细控制。
  • 在两项挑战中均领先,自然度、风格相似性与歌手相似性全面胜出。
  • 适用于音乐生成、语音克隆等需要精准风格迁移的场景。

我们提出S²Voice,该系统在2025年歌唱语音转换挑战赛(SVCC 2025)的域内与零样本转换任务中均获冠军。基于强大的两阶段Vevo基线,S²Voice通过多项改进提升风格控制力与鲁棒性:首先,在自回归大语言模型中引入风格嵌入,采用FiLM式层归一化条件和风格感知交叉注意力以实现细粒度风格建模;其次,在流匹配变换器中加入全局说话人嵌入,提升音色相似性;第三,通过自动化网络采集、人声分离与转录精修流程构建大规模高质量歌唱语料库;最后,采用监督微调(SFT)与直接偏好优化(DPO)相结合的多阶段训练策略。主观听感测试验证了其优越性能:在任务1中风格相似性与歌手相似性领先,在任务2中自然度、风格相似性与歌手相似性全面领先。消融实验表明各项改进有效提升了风格保真度、音色保留与泛化能力。音频样例可访问:https://honee-w.github.io/SVC-Challenge-Demo/

原文摘要 · Abstract (English)

We present S$^2$Voice, the winning system of the Singing Voice Conversion Challenge (SVCC) 2025 for both the in-domain and zero-shot singing style conversion tracks. Built on the strong two-stage Vevo baseline, S$^2$Voice advances style control and robustness through several contributions. First, we integrate style embeddings into the autoregressive large language model (AR LLM) via a FiLM-style layer-norm conditioning and a style-aware cross-attention for enhanced fine-grained style modeling. Second, we introduce a global speaker embedding into the flow-matching transformer to improve timbre similarity. Third, we curate a large, high-quality singing corpus via an automated pipeline for web harvesting, vocal separation, and transcript refinement. Finally, we employ a multi-stage training strategy combining supervised fine-tuning (SFT) and direct preference optimization (DPO). Subjective listening tests confirm our system's superior performance: leading in style similarity and singer similarity for Task 1, and across naturalness, style similarity, and singer similarity for Task 2. Ablation studies demonstrate the effectiveness of our contributions in enhancing style fidelity, timbre preservation, and generalization. Audio samples are available~\footnote{https://honee-w.github.io/SVC-Challenge-Demo/}.

语音转换风格迁移歌唱合成自回归模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。