arXiv:2603.25035cs.AI2026-03

分析视觉语言模型压缩对内部机制和安全行为的影响

Mechanistically Interpreting Compression in Vision-Language Models

  • 用因果电路分析与特征对比研究剪枝和量化对模型内部的影响
  • 剪枝使特征旋转衰减,量化则提升特征对齐度但降低拒答率
  • 提出新基准VLMSafe-420,揭示压缩方式影响模型安全表现

压缩的视觉语言模型(VLMs)广泛用于降低内存和计算开销,适合实际部署。然而,压缩是否保留内部计算与安全行为仍存疑。本文通过因果电路分析和crosscoder特征对比,考察剪枝与量化对代表性VLMs内部结构的根本影响。结果表明,剪枝通常保持电路结构完整,但导致内部特征旋转与衰减;量化在更高层级修改电路,却使幸存特征更对齐。基于此,我们提出VLMSafe-420,一个新基准,将有害输入与其匹配的良性反事实样本配对,覆盖多种安全类别。实验显示,剪枝显著降低真实拒答行为,说明压缩策略具有安全影响。

原文摘要 · Abstract (English)

Compressed vision-language models (VLMs) are widely used to reduce memory and compute costs, making them a suitable choice for real-world deployment. However, compressing these models raises concerns about whether internal computations and safety behaviors are preserved. In this work, we use causal circuit analysis and crosscoder-based feature comparisons to examine how pruning and quantization fundamentally change the internals across representative VLMs. We observe that pruning generally keeps circuit structure intact but rotates and attenuates internal features, while quantization modifies the circuits at a higher level yet leaves the surviving features better aligned. Leveraging this insight, we also introduce VLMSafe-420, a novel benchmark that pairs harmful inputs with matched benign counterfactuals across various safety categories. Our findings show that pruning causes a sharp drop in genuine refusal behavior, suggesting that the choice of compression has safety implications.

模型压缩视觉语言模型安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。