让音乐生成模型可精准调控音高和时长,无需重新训练。
Closing the Loop: PID Feedback Control for Interpretable Activation Steering in Symbolic Music Generation

- 通过分析隐藏层激活,找到控制音高和时长的潜在方向。
- 使用几何去耦技术,实现多属性独立调控且不互相干扰。
- 适合需要精确控制音乐属性的研究者与创作者使用。
基于Transformer的架构显著提升了复杂符号序列的生成能力,但在细粒度、可解释的离散信号属性控制方面仍存在明显差距。本文研究了多轨音乐Transformer(MMT)的机制可解释性,提出一种无需重训练即可在推理阶段实现确定性属性调制的框架。利用均值差异(DiffMean)方法,我们在残差流中识别出音高和时长等信号属性的潜在方向。验证了该领域中的线性表示假设,发现调制幅度与属性变化高度相关。为解决多属性调制中的特征纠缠问题,引入基于格拉姆-施密特正交化的双路径调制框架。实验表明,该几何解耦方法相比直接向量叠加,显著降低了概念干扰和信号退化,即使在强自回归条件约束下也能实现独立确定性控制。
原文摘要 · Abstract (English)
Transformer-based architectures have significantly advanced the generation of complex symbolic sequences, yet a significant gap remains in achieving fine-grained, interpretable control over discrete signal attributes. This paper investigates the mechanistic interpretability of the Multitrack Music Transformer (MMT) and proposes a framework for deterministic attribute modulation without retraining to bridge this gap via inference-time activation steering. Utilizing the Difference-in-Means (DiffMean) methodology, we isolate latent directions for signal attributes, specifically Pitch and Duration, within the residual stream. We validate the Linear Representation Hypothesis in this domain, achieving high correlation between steering magnitude and attribute shift. To address the inherent feature entanglement in multi-attribute steering, we introduce a Dual Steering framework utilizing Gram-Schmidt Orthogonalization. Experimental results demonstrate that this geometric decoupling reduces conceptual interference and signal degradation compared to naive vector addition, enabling independent deterministic control even against strong autoregressive conditioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。