arXiv:2411.14773cs.SDcs.AI2024-11被引 1

用类脑神经网络模拟人类对音乐调式的认知,生成有调性特征的乐曲。

Mode-conditioned music learning and composition: a spiking neural network inspired by neuroscience and psychology

  • 基于脑区结构和心理机制设计多子系统神经网络,模拟人类对调式感知。
  • 生成乐曲具备指定调式与键的特征,且旋律多样而自然。
  • 模型结构与音乐心理学经典理论高度吻合,适合音乐生成与认知研究者。

音乐调式是构建音高组织框架并决定和声关系的核心要素。以往方法常采用简单僵化的对齐策略,忽视调式的多样性。然而,人类具备感知不同调式与调性的认知机制。本文提出一种受神经科学与心理学启发的脉冲神经网络,用于表征音乐调式与调性,并生成包含调性特征的乐曲。具体贡献包括:1)模型设计融合多个受大脑对应区域结构与功能启发的协同子系统;2)引入神经回路演化学习机制,使网络能学习并生成与调式相关的音乐特征,反映人类音乐感知的认知过程;3)实验结果表明,该模型的连接结构与音乐心理学领域重要模型Krumhansl-Schmuckler模型高度相似;4)生成乐曲在定量评估中展现出良好的调性特征与旋律适应性,能够生成多样化且具有音乐性的内容。本研究结合神经科学、心理学与音乐理论,推动人工智能音乐生成与人类认知的融合。

原文摘要 · Abstract (English)

Musical mode is one of the most critical element that establishes the framework of pitch organization and determines the harmonic relationships. Previous works often use the simplistic and rigid alignment method, and overlook the diversity of modes. However, in contrast to AI models, humans possess cognitive mechanisms for perceiving the various modes and keys. In this paper, we propose a spiking neural network inspired by brain mechanisms and psychological theories to represent musical modes and keys, ultimately generating musical pieces that incorporate tonality features. Specifically, the contributions are detailed as follows: 1) The model is designed with multiple collaborated subsystems inspired by the structures and functions of corresponding brain regions; 2)We incorporate mechanisms for neural circuit evolutionary learning that enable the network to learn and generate mode-related features in music, reflecting the cognitive processes involved in human music perception. 3)The results demonstrate that the proposed model shows a connection framework closely similar to the Krumhansl-Schmuckler model, which is one of the most significant key perception models in the music psychology domain. 4) Experiments show that the model can generate music pieces with characteristics of the given modes and keys. Additionally, the quantitative assessments of generated pieces reveals that the generating music pieces have both tonality characteristics and the melodic adaptability needed to generate diverse and musical content. By combining insights from neuroscience, psychology, and music theory with advanced neural network architectures, our research aims to create a system that not only learns and generates music but also bridges the gap between human cognition and artificial intelligence.

音乐生成脉冲神经网络调性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。