arXiv:2602.11910cs.SDcs.LG2026-02被引 1

通过激活调控实现对音频扩散模型的精细音乐属性控制

TADA! Tuning Audio Diffusion Models through Activation Steering

  • 发现音频扩散模型存在共享语义瓶颈,少数注意力层控制多种音乐概念
  • 激活调控在音乐属性调节上超越提示、分数空间和权重空间干预方法
  • 适用于需要精确控制乐器、人声或风格的音乐生成场景

音频扩散模型能从文本生成高保真音乐,但对特定音乐属性的细粒度控制仍具挑战,因其内部高层概念表示机制不清晰。本文通过激活补丁技术揭示,近期音频扩散架构存在语义瓶颈:少量连续注意力层共享控制不同音乐概念,如特定乐器、人声或流派。在此基础上,系统评估了多种调控范式,对比了激活调控与提示级、得分空间及权重空间干预的效果,并分析调控机制与干预位置的交互。新基准测试结合大规模用户研究显示,局部激活调控在音频概念调制上达到新最优性能。

原文摘要 · Abstract (English)

Audio diffusion models can synthesize high-fidelity music from text, yet achieving fine-grained control over specific musical attributes remains challenging, as their internal mechanisms for representing high-level concepts are poorly understood. In this work, we use activation patching to demonstrate that recent audio diffusion architectures exhibit a semantic bottleneck, where a small, shared subset of consecutive attention layers controls distinct musical concepts, such as the presence of specific instruments, vocals, or genres. Building on this, we systematically evaluate a broad spectrum of steering paradigms, comparing activation steering against prompt-level, score-space, and weight-space interventions, analyzing the interaction between the steering mechanism and the intervention site. Our new benchmark, supported by an extensive user study, demonstrates that localized activation steering establishes a new state-of-the-art in audio concept modulation.

音频生成扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。