arXiv:2510.23802cs.LGcs.SD2025-10中稿 · NeurIPS被引 8

用稀疏自编码器让音频生成模型的隐空间变得可解释

Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders

  • 在音频隐空间上训练稀疏自编码器,提取可读的声学特征
  • 建立特征与音高、响度、音色等属性的线性映射关系
  • 适用于音乐生成模型分析,帮助理解声音合成过程

尽管稀疏自编码器(SAEs)在语言模型中成功提取出可解释特征,但将其应用于音频生成仍面临挑战:音频信号本身密集,压缩会模糊语义,且自动特征表征能力有限。本文提出一种框架,通过将音频生成模型的隐表示映射到人类可理解的声学概念,实现对音频生成过程的解释。我们在音频自编码器的隐空间上训练SAE,再学习从SAE特征到离散声学属性(音高、振幅、音色)的线性映射。该方法支持可控编辑与过程分析,揭示了合成过程中声学属性的演化机制。我们在连续(DiffRhythm-VAE)和离散(EnCodec、WavTokenizer)音频隐空间上验证了该方法,并以先进文本到音乐模型DiffRhythm为例,展示了音高、音色和响度在整个生成过程中的演变。虽然本研究仅针对音频模态,但该框架可拓展至视觉生成模型的可解释分析。

原文摘要 · Abstract (English)

While sparse autoencoders (SAEs) successfully extract interpretable features from language models, applying them to audio generation faces unique challenges: audio's dense nature requires compression that obscures semantic meaning, and automatic feature characterization remains limited. We propose a framework for interpreting audio generative models by mapping their latent representations to human-interpretable acoustic concepts. We train SAEs on audio autoencoder latents, then learn linear mappings from SAE features to discretized acoustic properties (pitch, amplitude, and timbre). This enables both controllable manipulation and analysis of the AI music generation process, revealing how acoustic properties emerge during synthesis. We validate our approach on continuous (DiffRhythm-VAE) and discrete (EnCodec, WavTokenizer) audio latent spaces, and analyze DiffRhythm, a state-of-the-art text-to-music model, to demonstrate how pitch, timbre, and loudness evolve throughout generation. While our work is only done on audio modality, our framework can be extended to interpretable analysis of visual latent space generation models.

音频生成可解释性稀疏编码隐空间分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。