arXiv:2504.02586cs.SDcs.AI2025-04被引 1

对比四种AI音乐生成方法,GPT3效果最佳,新节奏理论表现突出。

Deep learning for music generation. Four approaches and their comparative evaluation

  • 用视觉Transformer、ChatSonification、Schillinger理论与GPT3生成音乐
  • GPT3生成旋律最悦耳,新方法优于传统声学转化法
  • 适合音乐创作、跨模态生成研究者参考

本文介绍四种不同的AI音乐生成算法,并从生成音乐的审美质量及特定应用场景适配性两方面进行比较。第一种使用微调的视觉Transformer作为语言模型生成旋律;第二种结合聊天声学化与经典Transformer(此前已有研究);第三种融合施林格节奏理论与经典Transformer;第四种采用OpenAI提供的GPT3 Transformer。对生成旋律的对比分析显示,各方法间存在显著差异。在审美价值方面,GPT3生成的旋律最为悦耳;新引入的施林格方法生成的音乐也优于以往的声学化方法。

原文摘要 · Abstract (English)

This paper introduces four different artificial intelligence algorithms for music generation and aims to compare these methods not only based on the aesthetic quality of the generated music but also on their suitability for specific applications. The first set of melodies is produced by a slightly modified visual transformer neural network that is used as a language model. The second set of melodies is generated by combining chat sonification with a classic transformer neural network (the same method of music generation is presented in a previous research), the third set of melodies is generated by combining the Schillinger rhythm theory together with a classic transformer neural network, and the fourth set of melodies is generated using GPT3 transformer provided by OpenAI. A comparative analysis is performed on the melodies generated by these approaches and the results indicate that significant differences can be observed between them and regarding the aesthetic value of them, GPT3 produced the most pleasing melodies, and the newly introduced Schillinger method proved to generate better sounding music than previous sonification methods.

音乐生成TransformerGPT3节奏理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。