用和弦控制生成歌曲,提升音乐表现力与精确度。
An End-to-End Approach for Chord-Conditioned Song Generation
- 引入和弦信息作为生成条件,通过动态加权注意力融合
- 在多个指标上优于现有方法,尤其提升旋律与伴奏契合度
- 适合需要精准控制音乐风格的作曲与编曲场景
歌曲生成任务旨在根据给定歌词合成包含人声与伴奏的完整音乐。现有方法如Jukebox虽已探索该任务,但生成过程中对音乐表现力的控制能力受限,常导致伴奏与旋律不协调。为解决此问题,本文引入音乐创作中的核心概念——和弦,作为伴奏基础并为旋律提供和声支持。针对自动和弦提取不准确的问题,提出一种增强型交叉注意力机制,结合动态权重序列,有效融合和弦信息,减少生成过程中的帧级错误。基于该机制,构建了新的端到端模型——和弦条件歌曲生成器(Chord-Conditioned Song Generator, CSG)。实验结果表明,该方法在音乐表现力和生成控制精度方面均优于现有方法。
原文摘要 · Abstract (English)
The Song Generation task aims to synthesize music composed of vocals and accompaniment from given lyrics. While the existing method, Jukebox, has explored this task, its constrained control over the generations often leads to deficiency in music performance. To mitigate the issue, we introduce an important concept from music composition, namely chords, to song generation networks. Chords form the foundation of accompaniment and provide vocal melody with associated harmony. Given the inaccuracy of automatic chord extractors, we devise a robust cross-attention mechanism augmented with dynamic weight sequence to integrate extracted chord information into song generations and reduce frame-level flaws, and propose a novel model termed Chord-Conditioned Song Generator (CSG) based on it. Experimental evidence demonstrates our proposed method outperforms other approaches in terms of musical performance and control precision of generated songs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。