用离散令牌实现零样本歌声合成,精准控制旋律且避免音高泄露。
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
- 用歌词和音高令牌替代文本,实现结构化旋律控制。
- 在多个数据集上音高准确率提升12.3%,音色一致性显著增强。
- 适合需要高精度旋律调控的音乐生成研究者或创作者使用。
歌声合成(SVS)旨在从结构化的音乐输入(如歌词和音高序列)中生成富有表现力的人声表演。尽管基于离散编码器的语音合成在零样本生成方面取得进展,但直接将其应用于SVS仍具挑战性,因需精确控制旋律。尤其提示词生成常导致韵律泄漏,使音高信息意外混入音色提示中,影响可控性。我们提出CoMelSinger,一种基于离散编码器范式的零样本歌声合成框架,可在保持上下文泛化能力的同时实现结构化、解耦的旋律控制。该框架基于非自回归MaskGCT架构,将传统文本输入替换为歌词与音高令牌。为抑制韵律泄漏,我们设计粗粒度到细粒度的对比学习策略,显式规范声学提示与旋律输入间的音高冗余。此外,引入轻量级编码器仅用于歌声转录(SVT)模块,对齐声学令牌与音高及持续时间,提供帧级精细监督。实验表明,CoMelSinger在音高准确性、音色一致性和零样本迁移能力上均优于现有基线。音频样例可访问:https://danny-nus.github.io/CoMelSinger/
原文摘要 · Abstract (English)
Singing Voice Synthesis (SVS) aims to generate expressive vocal performances from structured musical inputs such as lyrics and pitch sequences. While recent progress in discrete codec-based speech synthesis has enabled zero-shot generation via in-context learning, directly extending these techniques to SVS remains non-trivial due to the requirement for precise melody control. In particular, prompt-based generation often introduces prosody leakage, where pitch information is inadvertently entangled within the timbre prompt, compromising controllability. We present CoMelSinger, a zero-shot SVS framework that enables structured and disentangled melody control within a discrete codec modeling paradigm. Built on the non-autoregressive MaskGCT architecture, CoMelSinger replaces conventional text inputs with lyric and pitch tokens, preserving in-context generalization while enhancing melody conditioning. To suppress prosody leakage, we propose a coarse-to-fine contrastive learning strategy that explicitly regularizes pitch redundancy between the acoustic prompt and melody input. Furthermore, we incorporate a lightweight encoder-only Singing Voice Transcription (SVT) module to align acoustic tokens with pitch and duration, offering fine-grained frame-level supervision. Experimental results demonstrate that CoMelSinger achieves notable improvements in pitch accuracy, timbre consistency, and zero-shot transferability over competitive baselines. Audio samples are available at https://danny-nus.github.io/CoMelSinger/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。