arXiv:2409.12638cs.SDcs.HC2024-09被引 2

用自然语言生成可调节的多轨长时音乐,支持任意拍号和风格。

M6(GPT)3: Generating Multitrack Modifiable Multi-Minute MIDI Music from Text using Genetic algorithms, Probabilistic methods and GPT Models in any Progression and Time Signature

  • 结合遗传算法与GPT模型,从文本生成多轨复杂乐曲结构。
  • 在情感参数驱动下自适应演化旋律,生成效果优于基线。
  • 适合音乐创作、游戏配乐等需要灵活生成长音乐的场景。

本文提出M6(GPT)3作曲系统,能够根据自然语言描述生成完整、多分钟且结构复杂的音乐作品,支持任意拍号,在MIDI领域实现多轨输出。系统采用自回归Transformer语言模型将文本提示映射为包含拍号、调式、和弦进行及情绪值(愉悦-唤醒)的JSON格式参数。基于这些参数生成伴奏、主旋律、低音、动机与打击乐轨道。针对旋律生成,提出基于音乐意义突变与正态分布适应性评估的遗传算法,其适应度函数受情绪参数与演奏风格影响动态调整。打击乐生成采用马尔可夫链等概率方法,支持任意拍号。通过人工与客观评估验证,本方法在多个音乐有意义指标上优于基线,为纯神经网络系统提供可行替代方案。

原文摘要 · Abstract (English)

This work introduces the M6(GPT)3 composer system, capable of generating complete, multi-minute musical compositions with complex structures in any time signature, in the MIDI domain from input descriptions in natural language. The system utilizes an autoregressive transformer language model to map natural language prompts to composition parameters in JSON format. The defined structure includes time signature, scales, chord progressions, and valence-arousal values, from which accompaniment, melody, bass, motif, and percussion tracks are created. We propose a genetic algorithm for the generation of melodic elements. The algorithm incorporates mutations with musical significance and a fitness function based on normal distribution and predefined musical feature values. The values adaptively evolve, influenced by emotional parameters and distinct playing styles. The system for generating percussion in any time signature utilises probabilistic methods, including Markov chains. Through both human and objective evaluations, we demonstrate that our music generation approach outperforms baselines on specific, musically meaningful metrics, offering a viable alternative to purely neural network-based systems.

音乐生成GPT遗传算法多轨音乐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。