arXiv:2505.18186cs.SDcs.LG2025-05被引 12

用稀疏自编码器挖掘音乐生成模型中的可解释概念。

Discovering and Steering Interpretable Concepts in Large Generative Music Models

  • 通过稀疏自编码器从Transformer残差流中提取可解释特征。
  • 发现既有传统音乐理论中的概念,也有未被命名的新模式。
  • 可用来引导模型生成,提升可控性与透明度。

神经网络生成音乐的高保真度带来科学机遇:这些系统仅通过统计学习就可能内化内容结构的隐含理论。当内部表征与传统概念(如和弦进行)对齐时,揭示了此类范畴如何从统计规律中涌现;当偏离时,则暴露现有框架的局限性及被忽视但具解释力的模式。本文聚焦自回归音乐生成器,提出一种基于稀疏自编码器(SAEs)的方法,从Transformer模型的残差流中提取可解释特征,并通过自动化标注与验证流程实现方法的可扩展性与可评估性。结果发现既有熟悉的音乐概念,也存在结构一致但缺乏理论或语言对应的新模式。进一步证明这些概念可用于引导模型生成。本工作不仅增强模型透明度,还提供了一种实证工具,用于发现传统分析与合成方法难以捕捉的组织原则。

原文摘要 · Abstract (English)

The fidelity with which neural networks can now generate content such as music presents a scientific opportunity: these systems appear to have learned implicit theories of such content's structure through statistical learning alone. This offers a potentially new lens on theories of human-generated media. When internal representations align with traditional constructs (e.g. chord progressions in music), they show how such categories can emerge from statistical regularities; when they diverge, they expose limits of existing frameworks and patterns we may have overlooked but that nonetheless carry explanatory power. In this paper, focusing on autoregressive music generators, we introduce a method for discovering interpretable concepts using sparse autoencoders (SAEs), extracting interpretable features from the residual stream of a transformer model. We make this approach scalable and evaluable using automated labeling and validation pipelines. Our results reveal both familiar musical concepts and coherent but uncodified patterns lacking clear counterparts in theory or language. As an extension, we show such concepts can be used to steer model generations. Beyond improving model transparency, our work provides an empirical tool for uncovering organizing principles that have eluded traditional methods of analysis and synthesis.

音乐生成可解释性自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。