MPEcho通过显式音素条件控制,提升翻唱生成的歌词准确性。
MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation

- 引入音素编码器与长度调节器,实现音素级精确控制
- 音素错误率(PER)显著降低,歌词对齐更准确
- 适合需要高歌词保真度的音乐生成任务
翻唱歌曲生成(CSG)需在保留参考歌曲旋律和语言内容的同时重构其余音乐成分。当前先进模型SongEcho采用基频(F₀)序列和有声/无声(V/UV)标签进行条件约束;然而,V/UV标签中隐含的语言信息无法保证歌词准确性,导致音素错误率(PER)较高。受歌声合成(SVS)启发,我们提出MPEcho,将音素编码器与长度调节器(LR)融入SongEcho框架。通过提供显式音素级条件输入与精确时间边界,MPEcho显著降低PER。为此,我们开发了Phonsa——一个基于Whisper的自动转录模型,可为演唱语音生成高精度音素级标注,缓解高质量音频-音素配对数据稀缺问题。实验验证了Phonsa在对齐上的有效性及MPEcho在端到端翻唱生成中的优越表现。音频样本、代码与权重可通过https://lonian6.github.io/MPEcho.github.io/获取。
原文摘要 · Abstract (English)
Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components. The state-of-the-art model SongEcho utilizes $F_0$ sequences and voiced/unvoiced (V/UV) tags for conditioning; however, implicit linguistic information from V/UV tags cannot guarantee lyric accuracy, leading to a high phoneme error rate (PER). Inspired by singing voice synthesis (SVS), we propose MPEcho, which integrates a phoneme encoder and a length regulator (LR) into the SongEcho framework. By providing explicit phoneme-level conditioning and precise temporal boundaries, MPEcho significantly reduces PER. To enable this, we developed Phonsa, a Whisper-based automatic transcription model that provides high-precision phoneme-level annotations for singing voices, overcoming the scarcity of high-quality audio-phoneme pairs. Experimental results validate the effectiveness of Phonsa for alignment and MPEcho for end-to-end CSG. The audio samples, code and weights can be accessed from https://lonian6.github.io/MPEcho.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。