arXiv:2509.24603cs.SDcs.CV2025-09

无监督学习发现音乐中的基础模式,助力生成与分析。

Discovering "Words" in Music: Unsupervised Learning of Compositional Sparse Code for Symbolic Music

  • 通过两阶段EM框架构建音乐词典并重建乐谱
  • 在人类标注上达到0.61的交并比(IoU)
  • 适合音乐生成、风格迁移及音乐学研究

本文提出一种无监督机器学习算法,从符号化音乐数据中识别重复出现的结构单元——称为“音乐词”。这些模式是音乐结构的基础,反映作曲的认知过程。由于音乐解释存在固有的语义模糊性,提取这些模式仍具挑战。我们将音乐词发现任务建模为统计优化问题,提出基于两阶段期望最大化(EM)的学习框架:1. 构建音乐词典;2. 重构音乐数据。在与人类专家标注对比时,该算法取得0.61的交并比(IoU)得分。研究发现,最小化编码长度能有效缓解语义模糊,表明人类对编码系统的优化塑造了音乐语义。该方法使计算机能够从音乐数据中提取“基本构建块”,支持结构分析与稀疏编码。主要应用包括:在人工智能音乐中,支撑音乐生成、分类、风格迁移与即兴创作等下游任务;在音乐学中,提供分析作曲模式的工具,并揭示不同音乐风格与作曲家间最小编码原则的共性。

原文摘要 · Abstract (English)

This paper presents an unsupervised machine learning algorithm that identifies recurring patterns -- referred to as ``music-words'' -- from symbolic music data. These patterns are fundamental to musical structure and reflect the cognitive processes involved in composition. However, extracting these patterns remains challenging because of the inherent semantic ambiguity in musical interpretation. We formulate the task of music-word discovery as a statistical optimization problem and propose a two-stage Expectation-Maximization (EM)-based learning framework: 1. Developing a music-word dictionary; 2. Reconstructing the music data. When evaluated against human expert annotations, the algorithm achieved an Intersection over Union (IoU) score of 0.61. Our findings indicate that minimizing code length effectively addresses semantic ambiguity, suggesting that human optimization of encoding systems shapes musical semantics. This approach enables computers to extract ``basic building blocks'' from music data, facilitating structural analysis and sparse encoding. The method has two primary applications. First, in AI music, it supports downstream tasks such as music generation, classification, style transfer, and improvisation. Second, in musicology, it provides a tool for analyzing compositional patterns and offers insights into the principle of minimal encoding across diverse musical styles and composers.

音乐生成无监督学习符号音乐结构分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。