用视频唇动信息实现无需时长标注的高质量歌声合成
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
- 融合视频唇动特征,实现无须声学时长标注的歌声合成
- 在主观与客观评测中均达到当前最佳效果
- 适合语音合成、跨模态生成及影视配音领域研究者
现有歌声合成(SVS)模型主要依赖精细的音素级时长标注,限制了实际应用。这些方法忽略了视觉信息在时长预测中的互补作用。为此,我们提出 PerformSinger,首个结合演唱视频唇部动作作为视觉模态的多模态歌声合成框架,实现高质量的“无时长”歌声合成。PerformSinger 包含并行多分支多模态编码器、特征融合模块、时长与变分预测网络、梅尔频谱解码器和声码器。融合模块由适配器和融合块组成,在对齐语义空间中采用渐进式融合策略,生成高质量多模态特征表示,从而实现精准时长预测与高保真音频合成。为推动研究,我们构建并标注了一个包含同步视频流与精确音素级人工标注的新数据集。大量实验表明,该方法在主观与客观评价中均达到当前最优性能。代码与数据集将公开共享。
原文摘要 · Abstract (English)
Existing singing voice synthesis (SVS) models largely rely on fine-grained, phoneme-level durations, which limits their practical application. These methods overlook the complementary role of visual information in duration prediction.To address these issues, we propose PerformSinger, a pioneering multimodal SVS framework, which incorporates lip cues from video as a visual modality, enabling high-quality "duration-free" singing voice synthesis. PerformSinger comprises parallel multi-branch multimodal encoders, a feature fusion module, a duration and variational prediction network, a mel-spectrogram decoder and a vocoder. The fusion module, composed of adapter and fusion blocks, employs a progressive fusion strategy within an aligned semantic space to produce high-quality multimodal feature representations, thereby enabling accurate duration prediction and high-fidelity audio synthesis. To facilitate the research, we design, collect and annotate a novel SVS dataset involving synchronized video streams and precise phoneme-level manual annotations. Extensive experiments demonstrate the state-of-the-art performance of our proposal in both subjective and objective evaluations. The code and dataset will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。