arXiv:2608.08757cs.SDeess.AS2026-08中稿 · ISMIR 2026

用无监督方法精准定位音乐特征,实现更稳定的概念控制。

Steering dense music retrieval with open-vocabulary concept discovery

论文配图:Steering dense music retrieval with open-vocabulary concept discovery
图 1 · 摘自论文原文
  • 将概念归因转为音频空间的稀疏反演问题,不依赖文本匹配。
  • 在不重训练的前提下,使概念增强与抑制效果更强、保真度更高。
  • 适合需要精准控制音乐风格但无标注数据的研究者使用。

可控音乐检索允许用户找到如更氛围化、失真更少或无吉他等特定风格的音乐,同时保留原始查询的其他语义内容。稀疏自编码器(SAEs)是实现这种概念级控制的有力工具,但关键挑战在于:给定自由文本概念后,应编辑哪些稀疏特征?在共享多模态嵌入空间中,标准归因方法常选择与概念文字匹配但未对应实际音频示例的神经元,导致编辑效果弱或不稳定——当概念分布在多个神经元时遗漏相关特征,或因文本对齐而非音频结构被选中。为此,我们提出一种轻量级、无需训练的方法,恢复一组稀疏音频特征,其解码表示能重建目标概念,且保持音频空间几何一致性。该方法将概念归因重新定义为稀疏反演问题,而非基于文本侧的神经元排序启发式。无需成对音视频监督或SAE重训练。我们在可调控音乐检索任务中评估,结果表明恢复的支持更贴近承载概念的音频样本,在编辑强度与保留性能之间达成更优权衡,实现更精确的概念放大与抑制,同时降低保真度漂移。

原文摘要 · Abstract (English)

Controllable music retrieval lets users find music that is, for example, more ambient, less distorted, or without guitar while preserving the other semantic content of an original seed query. Sparse autoencoders (SAEs) are a promising interface for this kind of concept-level control, but a key problem remains: given a free-form text concept, which sparse features should be edited? In shared multimodal embedding spaces, standard attribution methods often select neurons that match the concept's wording but not the audio examples that express it. This leads to weak or unstable edits: relevant features are missed when concepts are distributed across neurons, while others are selected due to text alignment rather than audio-side structure. We address this with a lightweight, training-free method that recovers a sparse set of audio features whose decoded representation reconstructs the target concept while remaining consistent with audio-space geometry. This reframes concept attribution as a sparse inversion problem rather than a text-side neuron-ranking heuristic. The method requires neither paired audio-text supervision nor SAE retraining. We evaluate this approach in steerable music retrieval and show that the recovered supports align more closely with concept-bearing audio examples and achieve a stronger trade-off between edit strength and preservation than alignment baselines, enabling more precise concept amplification and suppression with reduced drift on preservation metrics.

音乐生成概念控制稀疏编码无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。