用轻量探针控制扩散模型音乐生成的音高,无需修改模型
Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes
- 训练小卷积探针从隐空间解码音高信息
- 引导生成使旋律连贯性提升2.4倍(显著优于基线)
- 适合想精准控制音乐音高的创作者使用
近期可控音乐生成多聚焦自回归模型,扩散模型研究相对不足。本文提出一种轻量级方法,用于控制基于稳定音频开放模型(Stable Audio Open)生成音乐的音高内容。通过配对音频与MIDI数据,训练一个约12.5万参数的小型卷积探针,从变分自编码器隐空间中解码帧级音高类激活。推理时,冻结的探针作为可微损失函数:其关于去噪隐变量的梯度用于引导生成向用户指定的音高序列靠拢,无需重训练或修改基础模型架构。在9个文本提示和3个目标旋律共27次评估中,探针引导生成的旋律连贯性比无引导基线提高2.4倍(p < 1e-5,Wilcoxon符号秩检验),证明扩散模型隐空间中蕴含可恢复且可操控的音乐结构。
原文摘要 · Abstract (English)
Recent work on controllable music generation has focused on autoregressive models, leaving diffusion-based systems comparatively underexplored. We present a lightweight method for steering the pitch content of audio produced by Stable Audio Open, a latent diffusion model for music synthesis. A small convolutional probe containing approximately 125k parameters is trained to decode frame-level pitch-class activations from the model's variational autoencoder latent space, using paired audio and MIDI data. At inference time, the frozen probe serves as a differentiable loss function: its gradient with respect to the denoising latent is used to nudge generation toward a user-specified pitch-class sequence, requiring no retraining or architectural modification of the base model. Across 27 evaluation trials spanning 9 text prompts and 3 target melodies, probe-guided generation increases melodic coherence by 2.4x over the unguided baseline (p < 1e-5, Wilcoxon signed-rank test), demonstrating that musically meaningful structure is both recoverable and steerable in diffusion-based music latent spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。