arXiv:2607.22000cs.SD2026-07

用钢琴动作预测声音,让模型学会音乐的因果关系。

Music-JEPA: Learning a World Model of Sound from Action

论文配图:Music-JEPA: Learning a World Model of Sound from Action
图 1 · 摘自论文原文
  • 把音频当状态、乐谱当动作,用未来音频预测当前状态。
  • 在无交互环境下训练,仍能捕捉音符与声音的因果联系。
  • 适合做音乐分析与自动配器,尤其对作曲家风格识别有帮助。

联合嵌入预测架构(JEPA)通过预测潜在表征来学习世界模型,为自监督学习提供了新方向。尽管已有研究尝试将JEPA应用于音乐领域,但其如何自然支持音乐世界模型的构建仍不明确。本文提出基于钢琴声音的音乐世界模型学习方法,将音乐视为动作条件系统:音频作为状态,钢琴记谱作为动作输入。给定当前音频状态和动作,模型预测未来音频状态,模拟人类通过互动学习音乐声音的过程。模型在离线条件下使用配对的音频-钢琴记谱数据进行训练,无需环境交互。实验表明,所学模型能有效捕捉音乐动作与其产生声音之间的关系。学习到的表征可支持多种下游任务,包括节拍追踪、作曲家识别和调性估计,并可通过规划实现钢琴转录——即搜索能最好解释目标声音的动作序列。

原文摘要 · Abstract (English)

Joint Embedding Predictive Architectures (JEPA) have recently emerged as a paradigm for learning world models by predicting latent representations, offering a promising direction for self-supervised learning. While initial attempts have applied JEPA to the music domain, it remains unclear how such frameworks can naturally support the formation of a world model for music. In this work, we propose to learn a world model of piano sound using JEPA by framing music as an action-conditioned system: the audio is treated as the state, and the pianoroll as the instrument action. Given a current audio state and an action, the model predicts the resulting future audio state, mirroring how humans learn musical sound through interaction. The model is trained in a fully offline setting using paired audio-pianoroll data, without environment interaction. Experiments show that the learned model captures the relationships between musical actions and their resulting sound. The resulting representations support downstream tasks, including beat tracking, composer identification, and key estimation, and enable piano transcription via planning, by searching for actions that best explain a target sound.

音乐生成世界模型自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。