用语言感知的韵律风格引导,提升低资源场景下的宝莱坞风格歌声合成质量。
LAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning
- 引入语言感知嵌入与声线引导学习机制,增强歌词语义与声线表达关联。
- 在小规模印地语数据集上,生成歌声自然度和表现力显著优于现有最优模型。
- 适合对跨语言、低资源语音合成及音乐风格建模感兴趣的开发者和研究者。
近年来,基于扩散模型的歌声合成(SVS)取得显著进展。然而,在低资源场景下,仍难以捕捉声线特征、特定流派的音高变化以及语言相关特性。为此,我们提出LAPS-Diff,一种融合语言感知嵌入与声线引导学习机制的扩散模型,专为宝莱坞印地语演唱风格设计。我们构建了一个印地语SVS数据集,并利用预训练语言模型提取词级和音素级嵌入,以丰富歌词表征。同时,引入风格编码器和音高提取模型,计算风格与音高损失,捕捉合成歌声的自然性与表现力关键特征。此外,采用MERT和IndicWav2Vec模型提取音乐与上下文嵌入,作为条件先验进一步优化声学特征生成。客观与主观评估表明,相较于所考虑的SOTA模型,LAPS-Diff在受限数据集上显著提升了生成样本的质量。
原文摘要 · Abstract (English)
The field of Singing Voice Synthesis (SVS) has seen significant advancements in recent years due to the rapid progress of diffusion-based approaches. However, capturing vocal style, genre-specific pitch inflections, and language-dependent characteristics remains challenging, particularly in low-resource scenarios. To address this, we propose LAPS-Diff, a diffusion model integrated with language-aware embeddings and a vocal-style guided learning mechanism, specifically designed for Bollywood Hindi singing style. We curate a Hindi SVS dataset and leverage pre-trained language models to extract word and phone-level embeddings for an enriched lyrics representation. Additionally, we incorporated a style encoder and a pitch extraction model to compute style and pitch losses, capturing features essential to the naturalness and expressiveness of the synthesized singing, particularly in terms of vocal style and pitch variations. Furthermore, we utilize MERT and IndicWav2Vec models to extract musical and contextual embeddings, serving as conditional priors to refine the acoustic feature generation process further. Based on objective and subjective evaluations, we demonstrate that LAPS-Diff significantly improves the quality of the generated samples compared to the considered state-of-the-art (SOTA) model for our constrained dataset that is typical of the low resource scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。