arXiv:2503.19611cs.SDcs.AI2025-03被引 24

让AI先构思音乐结构再生成,提升创作连贯性与质量

Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation

  • 用音乐思维链预设整体结构,再逐段生成音频
  • 在MUSDB18数据集上主观评分超越现有模型30%以上
  • 支持音色参考和结构分析,适合音乐创作与编辑者

自回归(AR)模型在高保真音乐生成中表现卓越,但传统的逐词预测方式不符合人类作曲的创造过程,可能影响生成作品的音乐性。为此,我们提出MusiCoT——一种专为音乐生成设计的思维链(CoT)提示方法。MusiCoT使AR模型在生成音频前先规划整体音乐结构,从而增强作品的连贯性与创意性。通过利用对比语言-音频预训练(CLAP)模型,构建“音乐思维链”,实现无需人工标注数据的可扩展性。此外,MusiCoT支持对乐器编排等结构进行深入分析,并可接受变长音频作为风格参考,有效缓解内容复制问题。实验表明,MusiCoT在客观与主观评估中均持续优于现有方法,生成音乐质量媲美顶尖模型。样本详见https://MusiCoT.github.io/

原文摘要 · Abstract (English)

Autoregressive (AR) models have demonstrated impressive capabilities in generating high-fidelity music. However, the conventional next-token prediction paradigm in AR models does not align with the human creative process in music composition, potentially compromising the musicality of generated samples. To overcome this limitation, we introduce MusiCoT, a novel chain-of-thought (CoT) prompting technique tailored for music generation. MusiCoT empowers the AR model to first outline an overall music structure before generating audio tokens, thereby enhancing the coherence and creativity of the resulting compositions. By leveraging the contrastive language-audio pretraining (CLAP) model, we establish a chain of "musical thoughts", making MusiCoT scalable and independent of human-labeled data, in contrast to conventional CoT methods. Moreover, MusiCoT allows for in-depth analysis of music structure, such as instrumental arrangements, and supports music referencing -- accepting variable-length audio inputs as optional style references. This innovative approach effectively addresses copying issues, positioning MusiCoT as a vital practical method for music prompting. Our experimental results indicate that MusiCoT consistently achieves superior performance across both objective and subjective metrics, producing music quality that rivals state-of-the-art generation models. Our samples are available at https://MusiCoT.github.io/.

音乐生成思维链高保真音频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。