arXiv:2512.10805cs.LGcs.CV2025-12被引 5

提升大模型可解释性,让关键概念可控调节

Interpretable and Steerable Concept Bottleneck Sparse Autoencoders

  • 提出新度量方法,系统评估神经元的可解释性和可控性
  • 改进后模型在可解释性上提升32.1%,可控性提升14.5%
  • 适合需要精准控制生成内容的场景,如图像编辑

稀疏自编码器(SAEs)为大语言模型和多模态模型提供了统一的机制可解释性、概念发现与模型调控路径。然而,其实现依赖于学习到的特征兼具可解释性与可控性。为此,我们引入两种计算开销小的可解释性与可控性度量,对视觉语言模型中的SAE进行系统分析,发现:(i) 多数神经元可解释性或可控性低,或两者兼有,难以用于下游任务;(ii) 用户期望的概念常缺失于SAE中,限制了实际应用。针对此,我们提出概念瓶颈稀疏自编码器(CB-SAE)——一种后处理框架,通过剪枝低效神经元并注入轻量级概念瓶颈以对齐用户定义的概念集。实验表明,该方法在多模态模型与图像生成任务中,可解释性提升32.1%,可控性提升14.5%。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential requires learned features to be both interpretable and steerable. To that end, we introduce two new computationally inexpensive interpretability and steerability metrics for a systematic analysis of LVLM SAEs. This uncovers two observations; (i) a majority of SAE neurons exhibit either low interpretability or low steerability or both, rendering them ineffective for downstream use; and (ii) user-desired concepts are often absent in the SAE, thus limiting their practical utility. To address these limitations, we propose Concept Bottleneck Sparse Autoencoders (CB-SAE) - a novel post-hoc framework that prunes low-utility neurons and augments the latent space with a lightweight concept bottleneck aligned to a user-defined concept set. The resulting CB-SAE improves interpretability by +32.1% and steerability by +14.5% across LVLMs and image generation tasks.

可解释性概念瓶颈自编码器模型控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。