用稀疏自编码器拆解音频大模型的模糊神经激活,让模型决策可解释。
AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs
- 通过稀疏自编码器将多义神经元激活分解为单一语义特征。
- 识别出可命名的代表性音效片段,并经人工验证有效。
- 适用于高风险场景透明化部署,支持未来扩展至多语言与细粒度语音特征。
尽管在音频感知任务中表现优异,大型音频-语言模型(AudioLLMs)仍缺乏可解释性。主要原因是模型中单个神经元常对多个无关概念产生响应。本文提出首个面向AudioLLMs的机制可解释性框架,利用稀疏自编码器(SAEs)将多义激活分解为单义特征。该流程识别代表性音频片段,通过自动字幕生成赋予有意义名称,并结合人工评估与控制实验验证概念有效性。实验表明,AudioLLMs能编码结构化且可解释的特征,提升透明度与可控性。本工作为高风险领域可信部署奠定基础,并支持未来扩展至更大模型、多语言音频及更细粒度的副语言特征。项目地址:https://townim-faisal.github.io/AutoInterpret-AudioLLM/
原文摘要 · Abstract (English)
Despite strong performance in audio perception tasks, large audio-language models (AudioLLMs) remain opaque to interpretation. A major factor behind this lack of interpretability is that individual neurons in these models frequently activate in response to several unrelated concepts. We introduce the first mechanistic interpretability framework for AudioLLMs, leveraging sparse autoencoders (SAEs) to disentangle polysemantic activations into monosemantic features. Our pipeline identifies representative audio clips, assigns meaningful names via automated captioning, and validates concepts through human evaluation and steering. Experiments show that AudioLLMs encode structured and interpretable features, enhancing transparency and control. This work provides a foundation for trustworthy deployment in high-stakes domains and enables future extensions to larger models, multilingual audio, and more fine-grained paralinguistic features. Project URL: https://townim-faisal.github.io/AutoInterpret-AudioLLM/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。