探究音乐生成模型是否内含乐理知识,发现其能识别乐理概念。
Do Music Generation Models Encode Music Theory?
- 构建合成乐理数据集,系统测试模型对乐理概念的感知能力。
- 大模型和深层结构更易识别乐理特征,效果随规模与层级变化。
- 为理解生成模型内在机制提供新视角,适合音乐人工智能研究者。
音乐基础模型具备强大的音乐生成能力。人类作曲时会运用音符、音程、和弦进行及速度等乐理知识构建旋律与节奏。这些模型是否也如此?我们提出 SynTheory,一个包含节拍、拍号、音符、音程、调式、和弦及和弦进行等概念的合成 MIDI 与音频数据集。基于此,构建框架探测音乐基础模型(Jukebox 与 MusicGen)内部表征中乐理概念的可辨识性。结果表明,乐理概念确实存在于模型内部,且可检测程度随模型规模与网络层深度而异。
原文摘要 · Abstract (English)
Music foundation models possess impressive music generation capabilities. When people compose music, they may infuse their understanding of music into their work, by using notes and intervals to craft melodies, chords to build progressions, and tempo to create a rhythmic feel. To what extent is this true of music generation models? More specifically, are fundamental Western music theory concepts observable within the "inner workings" of these models? Recent work proposed leveraging latent audio representations from music generation models towards music information retrieval tasks (e.g. genre classification, emotion recognition), which suggests that high-level musical characteristics are encoded within these models. However, probing individual music theory concepts (e.g. tempo, pitch class, chord quality) remains under-explored. Thus, we introduce SynTheory, a synthetic MIDI and audio music theory dataset, consisting of tempos, time signatures, notes, intervals, scales, chords, and chord progressions concepts. We then propose a framework to probe for these music theory concepts in music foundation models (Jukebox and MusicGen) and assess how strongly they encode these concepts within their internal representations. Our findings suggest that music theory concepts are discernible within foundation models and that the degree to which they are detectable varies by model size and layer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。