首次从神经元层面解析通用音频模型,发现其能自动产生跨类别识别神经元。
What Do Neurons Listen To? A Neuron-level Dissection of a General-purpose Audio Model
- 通过分析条件激活模式,定位类特异性神经元
- 神经元对语音属性、音乐音高等有共享响应,覆盖新任务类别
- 这些神经元直接影响分类性能,具实际功能意义
本文从神经元层级出发,剖析通用音频自监督学习(SSL)模型的内部表征。尽管这类模型作为特征提取器表现优异,但其鲁棒泛化能力背后的机制仍不明确。基于机制可解释性框架,我们通过分析多样任务下的条件激活模式,识别并检验类特异性神经元。结果表明,SSL模型会自发形成覆盖新任务类别的类特异性神经元,这些神经元在不同语义类别和声学相似性(如语音属性、音乐音高)上表现出共享响应。我们还验证了这些神经元对分类性能具有功能性影响。据我们所知,这是首个系统性的通用音频SSL模型神经元级分析,为理解其内部表示提供了新视角。
原文摘要 · Abstract (English)
In this paper, we analyze the internal representations of a general-purpose audio self-supervised learning (SSL) model from a neuron-level perspective. Despite their strong empirical performance as feature extractors, the internal mechanisms underlying the robust generalization of SSL audio models remain unclear. Drawing on the framework of mechanistic interpretability, we identify and examine class-specific neurons by analyzing conditional activation patterns across diverse tasks. Our analysis reveals that SSL models foster the emergence of class-specific neurons that provide extensive coverage across novel task classes. These neurons exhibit shared responses across different semantic categories and acoustic similarities, such as speech attributes and musical pitch. We also confirm that these neurons have a functional impact on classification performance. To our knowledge, this is the first systematic neuron-level analysis of a general-purpose audio SSL model, providing new insights into its internal representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。