arXiv:2607.14537cs.SDcs.LG2026-07被引 1

通过自监督学习构建符号音乐的层次化表示,支持高质量生成与理解。

MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

论文配图:MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music
图 1 · 摘自论文原文
  • 结合时移等变性目标与Swin Transformer,学习音乐的分层结构特征。
  • 重建F1达0.995,生成音乐在音高和节奏密度上高度匹配输入。
  • 适合音乐生成、作曲辅助及音乐表征研究者使用。

音乐结构的丰富内在表示对机器辅助作曲等音乐理解任务至关重要,但针对符号音乐的自监督表示学习仍较匮乏,尤其缺乏对音乐层级多尺度特性的建模。本文提出MIDI-RAE-JEPA,结合音高与时间平移等变性目标,采用LeJEPA框架与Swin Transformer V2编码器,从钢琴卷轴图像中学习符号音乐的层次化表示。时间平移等变性目标促使模型内化时间关系。编码器仅通过自监督目标(包括掩码嵌入预测)训练,并利用SIGReg防止表征坍塌。基于冻结编码器嵌入的解码器实现0.995的重建F1;条件流匹配生成模型生成的音乐在音高范围和节奏密度上紧密匹配输入片段,不匹配条件则生成虽无关但音乐上合理的输出。所学表示在下游情绪分类任务中优于哈爾散射变换基线,且嵌入距离随音高与时间偏移量单调递增,验证了可测量的等变性。结果表明,基于等变性的自监督学习目标,配合足够精细的编码器容量,为语义丰富、可用于生成的符号音乐表示提供了可行路径。

原文摘要 · Abstract (English)

Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures. We present MIDI-RAE-JEPA, combining a pitch- and time-shift equivariance objective with LeJEPA and a Swin Transformer V2 encoder to learn such hierarchical representations of symbolic music encoded as piano roll images. The time-shift equivariance objective encourages the model to internalize temporal musical relationships. The encoder is trained purely on self-supervised objectives -- including a masked embedding predictor (MEP) -- with collapse prevented via SIGReg. A separate decoder trained on the frozen encoder embeddings achieves reconstruction F1 of 0.995, and a flow matching generative model conditioned on those embeddings produces generations that closely match the pitch register and rhythmic density of the conditioning excerpt, while mismatched conditioning yields unrelated but musically plausible output. Learned representations outperform a Haar scattering transform baseline on a downstream emotion classification task, and embedding distances increase monotonically with pitch and time shift magnitude, confirming measurable equivariance. These results suggest that equivariance-based SSL objectives, combined with sufficient fine-level encoder capacity, provide a viable path toward semantically rich, generatively useful representations of symbolic music.

音乐生成自监督学习符号音乐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。