通过自发现概念向量实现音乐语义编辑,结构保持更完整。
AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing

- 从内部表征中自监督提取可解释的概念向量,无需标注数据。
- 在扩散模型隐空间注入概念向量,结构适配器保证节奏旋律不变。
- 适合需要高保真结构保留的音乐创作与编辑场景。
可控音乐编辑需在修改高层属性的同时严格保持节奏与旋律结构。然而,语义与结构常纠缠:传统方法为提升编辑效果牺牲结构完整性,而结构适配器又抑制语义响应。本文提出 AnchorSteer 框架,通过结构锚定与自发现语义引导解耦该矛盾。该方法利用自监督重建目标探测内部表征,提取无标签、可解释的概念向量,实现属性隔离。编辑时,这些可移植、即插即用的概念向量被注入扩散模型隐空间,同时结构适配器确保一致性。提供无条件与有条件注入变体以平衡鲁棒性与语义强度。在 ZoME-Bench 和主观评测中,所提框架显著优于仅靠引导或仅靠锚定的基线方法,实现显著语义变换且保持高保真结构。
原文摘要 · Abstract (English)
Controllable music editing is to modify high-level attributes while strictly preserving rhythmic and melodic structures. However, this task is challenged by a semantic-structural entanglement: steering methods often degrade structure to achieve editing performance, while structural adaptors suppress semantic responsiveness. We propose AnchorSteer, a framework that disentangles this tension by coupling structural anchoring with self-discovered semantic steering. The proposed approach probes internal representations to extract interpretable, label-free concept vectors via a self-supervised reconstruction objective, isolating attributes without curated data. During editing, these portable, plug-and-play concept vectors are injected into diffusion hidden manifolds while a structural adaptor enforces consistency. Variants for unconditioned and conditioned injections are provided to balance robustness and semantic strength. Experiments on ZoME-Bench and subjective tests show that the proposed framework outperforms both steering-only and anchoring-only baselines, enabling significant semantic transformations with high-fidelity structural preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。