arXiv:2410.01450cs.CLcs.AI2024-10中稿 · O-COCOSDA 2024被引 2

用多智能体系统生成更贴合旋律的中文歌词

Agent-Driven Large Language Models for Mandarin Lyric Generation

  • 设计多智能体分工协作,分别控制押韵、音节数、旋律契合度
  • 在Mpop600数据集上验证,智能体组合生成歌词与旋律匹配度更高
  • 适合音乐生成、歌词创作方向的研究者和开发者

生成式大语言模型展现出强大的上下文学习能力,仅通过提示即可完成多种任务。以往的旋律-歌词研究受限于高质量对齐数据稀缺及创作风格标准模糊,多数工作聚焦通用主题或情绪,而当前语言模型已具备足够能力,该方向价值降低。在汉语这类声调语言中,旋律与声调共同影响音高轮廓,导致歌词与旋律契合度存在差异。本研究基于Mpop600数据集验证发现,词作者与作曲者在创作时会主动考虑歌词与旋律的契合性。为此,我们构建了一个多智能体系统,将旋律-歌词生成任务分解为押韵、音节数、歌词-旋律对齐和一致性四个子任务,由不同智能体分别负责。通过基于扩散模型的歌声合成器进行听感测试,评估不同智能体组合生成歌词的质量。

原文摘要 · Abstract (English)

Generative Large Language Models have shown impressive in-context learning abilities, performing well across various tasks with just a prompt. Previous melody-to-lyric research has been limited by scarce high-quality aligned data and unclear standard for creativeness. Most efforts focused on general themes or emotions, which are less valuable given current language model capabilities. In tonal contour languages like Mandarin, pitch contours are influenced by both melody and tone, leading to variations in lyric-melody fit. Our study, validated by the Mpop600 dataset, confirms that lyricists and melody writers consider this fit during their composition process. In this research, we developed a multi-agent system that decomposes the melody-to-lyric task into sub-tasks, with each agent controlling rhyme, syllable count, lyric-melody alignment, and consistency. Listening tests were conducted via a diffusion-based singing voice synthesizer to evaluate the quality of lyrics generated by different agent groups.

歌词生成多智能体语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。