arXiv:2410.05151eess.AScs.SD2024-10中稿 · publication at ICA…被引 20

用旋律和文本控制音乐生成,提升长段音乐编辑质量。

Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer

  • 引入ControlNet增强DiT模型,实现文本与旋律双控生成。
  • 采用top-k常量Q变换表示旋律,减少多音轨模糊性。
  • 渐进式掩码策略稳定训练,适合音乐创作与编辑场景。

尽管可控音乐生成与编辑取得进展,但受限于梅尔频谱表示和UNet结构,仍存在生成质量与长度不足的问题。为此,我们提出一种新方法:在扩散Transformer(DiT)基础上增加控制分支,结合ControlNet实现由文本和旋律提示驱动的长时序、可变长度音乐生成与编辑。为实现更精细的旋律控制,提出一种新型top-k常量Q变换表示作为旋律提示,相比传统音高表示(如chroma),在多轨或宽音域音乐中显著降低歧义。通过渐进式掩码旋律提示的课程学习策略,有效平衡文本与旋律的控制信号,提升训练稳定性。实验基于开源乐器录音数据,在文本到音乐生成与风格迁移任务上验证效果。结果表明,扩展StableAudio预训练模型后,本方法在旋律控制编辑方面表现优异,同时保持良好的文本生成性能,优于MusicGen基线模型在文本生成与旋律保真度上的表现。音频示例见 https://stable-audio-control.github.io。

原文摘要 · Abstract (English)

Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structures. To address these limitations, we propose a novel approach using a Diffusion Transformer (DiT) augmented with an additional control branch using ControlNet. This allows for long-form and variable-length music generation and editing controlled by text and melody prompts. For more precise and fine-grained melody control, we introduce a novel top-$k$ constant-Q Transform representation as the melody prompt, reducing ambiguity compared to previous representations (e.g., chroma), particularly for music with multiple tracks or a wide range of pitch values. To effectively balance the control signals from text and melody prompts, we adopt a curriculum learning strategy that progressively masks the melody prompt, resulting in a more stable training process. Experiments have been performed on text-to-music generation and music-style transfer tasks using open-source instrumental recording data. The results demonstrate that by extending StableAudio, a pre-trained text-controlled DiT model, our approach enables superior melody-controlled editing while retaining good text-to-music generation performance. These results outperform a strong MusicGen baseline in terms of both text-based generation and melody preservation for editing. Audio examples can be found at https://stable-audio-control.github.io.

音乐生成扩散模型控制生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。