用旋律先验提升歌声合成质量,实现高保真可控变调。
MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation

- 分步融合旋律信息:先训练离散旋律分支,再融合特征。
- 实验显示音高一致性显著提升,变调时音色失真小。
- 适合音乐生成、语音合成研究者,尤其关注歌声质量者。
神经音频编解码器是基于大语言模型的音频生成的基础编码器。尽管语义先验被广泛用于提升语言可懂度,但显式声学先验的整合仍缺乏探索,限制了频率敏感领域(如歌声)的合成保真度。为填补这一空白,我们提出 MeloCodec,一种有效融入旋律先验(歌唱中关键的声学信息)的新框架。针对显式先验直接融合导致的优化不稳定性,我们设计了「先编码后融合」范式,预先训练离散旋律分支以锁定结构,再进行特征融合。为进一步稳定该范式,提出两阶段训练策略,防止码本坍塌并确保收敛。实验表明,MeloCodec 在歌声表征上优于基线方法,显著提升音高一致性,并实现可控音高调整,同时保持极低的音色失真。
原文摘要 · Abstract (English)
Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integration of explicit acoustic priors remains underexplored, limiting synthesis fidelity in frequency-sensitive domains. To address this gap, we introduce MeloCodec, a novel framework designed to effectively incorporate melodic priors, a critical form of acoustic information for singing. To address the optimization instability typically caused by the direct fusion of such explicit priors, we propose a Tokenize-then-Fuse paradigm that pre-trains a discrete melodic branch to lock in structures before feature fusion. To robustly realize this paradigm, we further propose a two-stage training strategy that prevents codebook collapse and ensures stable convergence. Experiments show that MeloCodec outperforms baselines in singing voice representation, improving pitch consistency and enabling controllable pitch manipulation with minimal timbre degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。