arXiv:2511.08252cs.SDeess.AS2025-11AAAI被引 6

不训练模型,通过干预注意力图实现精准音乐编辑。

Melodia: Training-Free Music Editing Guided by Attention Probing in Diffusion Models

  • 仅修改特定层自注意力图,保留原曲结构
  • 用注意力仓库存储源音乐信息,无需文本描述
  • 新评测指标更准确衡量编辑效果

文本生成音乐技术快速发展,但现有音乐编辑方法在改变乐器、风格或情绪等属性时,常破坏原曲的旋律与节奏结构。本文深入分析基于扩散模型的AudioLDM 2中的注意力图,发现跨注意力图虽包含音乐特征细节,但干预后常无效;而自注意力图对保持原曲时间结构至关重要。据此提出Melodia:一种无需训练的方法,在去噪过程中选择性操纵特定层的自注意力图,并利用注意力仓库存储源音乐信息,实现对音乐特征的精确修改,同时完整保留原始结构,且无需提供源音乐的文本描述。此外,提出两种新评测指标。客观与主观实验表明,该方法在多个数据集上均显著提升文本契合度与结构完整性。本研究深化了对音乐生成模型内部机制的理解,为音乐创作提供了更强控制力。

原文摘要 · Abstract (English)

Text-to-music generation technology is progressing rapidly, creating new opportunities for musical composition and editing. However, existing music editing methods often fail to preserve the source music's temporal structure, including melody and rhythm, when altering particular attributes like instrument, genre, and mood. To address this challenge, this paper conducts an in-depth probing analysis on attention maps within AudioLDM 2, a diffusion-based model commonly used as the backbone for existing music editing methods. We reveal a key finding: cross-attention maps encompass details regarding distinct musical characteristics, and interventions on these maps frequently result in ineffective modifications. In contrast, self-attention maps are essential for preserving the temporal structure of the source music during its conversion into the target music. Building upon this understanding, we present Melodia, a training-free technique that selectively manipulates self-attention maps in particular layers during the denoising process and leverages an attention repository to store source music information, achieving accurate modification of musical characteristics while preserving the original structure without requiring textual descriptions of the source music. Additionally, we propose two novel metrics to better evaluate music editing methods. Both objective and subjective experiments demonstrate that our approach achieves superior results in terms of textual adherence and structural integrity across various datasets. This research enhances comprehension of internal mechanisms within music generation models and provides improved control for music creation.

音乐生成扩散模型注意力机制编辑技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。