arXiv:2606.25657cs.CVcs.AI2026-06被引 1

提出联合稀疏自编码器,让视觉语言模型的跨模态控制更精准可解释。

Steering Vision-Language Models with Joint Sparse Autoencoders

论文配图:Steering Vision-Language Models with Joint Sparse Autoencoders
图 1 · 摘自论文原文
  • 通过显式对齐约束,联合分解视觉与语言激活,提取共享特征。
  • 在中后层实现最佳添加式干预效果,抑制效果则各层稳定。
  • 适用于多种模型架构,为多模态特征分析提供可控方向。

稀疏自编码器(SAEs)在语言模型分析中表现良好,但应用于视觉语言模型(VLMs)时,常得到难以用于可控跨模态调控的表示。我们提出联合稀疏自编码器(JSAE),通过显式对齐约束,将序列池化后的视觉与语言激活联合分解为共享的、可解释的图像/标题级特征。应用于LLaVA,JSAE成功恢复了可识别概念(如食物、动物)的跨模态特征。通过双向干预(添加式调控与抑制),我们发现:在实验协议下,添加式调控在中至晚期(预输出层)达到峰值,两端衰减;而抑制得分在所有探测层内波动接近统计噪声水平。在三种VLM上——LLaVA-v1.6-Mistral-7B、Llama3-LLaVA-8B 和基于MoE的Qwen3-VL-30B——均观察到类似层定位效应。结果表明,显式对齐的稀疏表示相比无约束方法,在可识别层范围内支持更可控的干预分析。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representations that are difficult to use as controllable cross-modal steering directions. We introduce the Joint Sparse Autoencoder (JSAE), which uses an explicit alignment constraint to jointly factorize sequence-pooled vision and language activations into shared, interpretable image/caption-level features. Applied to LLaVA, JSAE recovers cross-modal features for recognizable concepts (e.g., food and animals). Through bidirectional interventions (additive steering and suppression), we observe a layer-dependent asymmetry under our protocol: additive steering peaks at mid-to-late (pre-output) layers and weakens at both ends, whereas suppression scores remain within a comparable range across all probed layers within statistical noise. Experiments on three VLMs, namely LLaVA-v1.6-Mistral-7B, Llama3-LLaVA-8B, and the MoE-based Qwen3-VL-30B, show related layer-localized effects across architectures. Together, these results suggest that explicitly aligned sparse representations support more controllable intervention-based analysis of multimodal features, within an identifiable layer range, than the unconstrained alternatives tested here.

视觉语言模型稀疏编码可控生成跨模态分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。