arXiv:2506.01247cs.CVcs.AI2025-06被引 1

用稀疏编码器无标签操控视觉模型,提升零样本分类准确率

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering

  • 训练稀疏编码器提取图像特征,测试时放大活跃特征生成可解释控制向量
  • 在9个数据集上提升零样本准确率,最高增益+4.12%,推理开销低于0.1%
  • 发现重建重要特征未必对下游任务有用,揭示重构与任务显著性差异

稀疏自编码器(SAE)常用于解析基础模型,但其作为可操作干预空间的作用在视觉领域仍不明确。本文提出无需标签的视觉控制方法(VS2),在冻结的CLIP图像编码器激活值上训练顶-k SAE,测试时通过增强输入的活跃稀疏特征并解码变化,构建可解释的控制向量。该过程可形式化为质心-偏差控制:每个输入沿其与SAE学习质心的偏差方向移动,残差项由每样本重建误差(FVU)精确控制,据此建立基于FVU的残差界,并设计可靠性门控机制——当重建不可靠时退回到零样本CLIP。使用目标域未标注激活值训练的SAE,VS2在九个图像分类数据集上实现零样本准确率提升,最高达+4.12%,额外推理计算不足0.1%。进一步的上限研究VS2++表明,选择性放大稀疏特征可带来高达+21.44%的增益,揭示了重建显著性与下游任务显著性之间的差距。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) are increasingly used to interpret foundation models, but their role as an actionable intervention space remains less understood, especially in vision. We study whether sparse visual features can be used not only for post-hoc analysis, but also to steer frozen vision-language models. We introduce Visual Sparse Steering (VS2), a label-free method that trains a top-$k$ SAE on unlabeled activations from a frozen CLIP image encoder and, at test time, constructs an interpretable steering vector by amplifying the input's active sparse features and decoding the induced change. We show that this procedure admits a closed-form decomposition as centroid-deviation steering: each input is moved along its deviation from the SAE-learned centroid. The residual term is controlled exactly by the SAE's per-sample reconstruction error, measured by FVU, yielding an FVU-based residual bound and motivating a reliability gate that falls back to zero-shot CLIP when SAE reconstruction is unreliable. With target-domain SAEs trained on unlabeled CLIP image-encoder activations, VS2 improves zero-shot accuracy across nine image-classification datasets, achieving gains up to $+4.12\%$ with less than $0.1\%$ additional inference compute. Finally, a controlled upper-bound study, VS2++, shows that selective amplification of sparse features can yield gains up to $+21.44\%$, exposing a reconstruction-vs-task saliency gap: features salient for reconstruction need not align with features useful for downstream prediction.

稀疏编码视觉控制零样本可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。