arXiv:2603.25075cs.AI2026-03

发现视觉语言模型中稀疏特征电路不具模块性,干预易引发意外漂移。

Sparse Visual Thought Circuits in Vision-Language Models

  • 在多模态模型中定位稀疏特征电路,测试其可组合性
  • 联合干预导致预测漂移,准确率显著下降
  • 提供可复现的因果分析框架,适合模型可控性研究者

稀疏自编码器(SAEs)提升多模态模型可解释性,但其特征是否构成模块化、可组合的推理单元尚不明确。我们验证该模块性假设,发现其常失效:对任务选择性特征集进行干预可适度提升推理准确率,而对两个此类集合的并集进行干预则会引发显著输出漂移(预测大幅偏离)且降低准确率,即使在范数匹配扰动下亦如此。这种非模块化干扰与内部路径共享一致,特征并集放大激活偏移。我们构建了可复现的因果流程,用于定位和测试 Qwen3-VL-8B 中的稀疏视觉思维电路。在包含七类任务、三难度等级的合成基准上,线性探测识别出中间解码层为任务类型信息所在位置。在此层训练 SAE,通过显式规则构建任务选择性特征集,并在推理时进行缩放与消融实验,量化准确率与漂移程度。研究结果经自助抽样、置换控制验证,且在多个 VLM 家族与五个不同数据集上重现,明确了 SAE 特征可组合性的边界,提供了更可靠的 VLM 控制诊断框架。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) improve interpretability in multimodal models, but it remains unclear whether SAE features form modular, composable units for reasoning-an assumption underlying many intervention-based steering methods. We test this modularity hypothesis and find it often fails: intervening on a task-selective feature set can modestly improve reasoning accuracy, while intervening on the union of two such sets reliably induces output drift (large unintended changes in predictions) and degrades accuracy, even under norm-matched perturbations. This non modular circuit interference is consistent with shared internal pathways where feature unions amplify activation shifts. We develop a reproducible causal pipeline to localize and test these sparse visual thought circuits in Qwen3-VL-8B. On a controlled synthetic benchmark with seven task types and three difficulty levels, linear probes identify a mid decoder locus for task type information. We train SAEs at this layer, construct task-selective sets via an explicit rule, and perform inference time scaling and ablation while quantifying accuracy and drift. Our findings-validated with bootstrapped subsamples and permutation controls, and replicated across multiple VLM families and five diverse datasets clarify the boundaries of SAE feature composability and provide a rigorous diagnostic framework for more reliable VLM control.

可解释性视觉语言模型稀疏编码因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。