提出首个评估视觉模型概念干预效果的基准,揭示正交化对鲁棒性的影响。
SwordBench: Evaluating Orthogonality of Steering Image Representations

- 设计跨模型、多任务的图像表征干预评估基准SwordBench
- 发现线性SVM虽正交性好但存在显著副作用,稀疏自编码器表现更优
- 适用于关注模型可解释性与安全性的研究者
在推理阶段干预模型表征以修正预测,对人工智能可解释性与安全性至关重要,但现有评估方法仅限于模糊的语言建模任务。为填补这一空白,我们提出SwordBench,一个针对多种骨干网络和概念移除任务的视觉模型表征干预基准。除了统一的评估套件,我们还引入新评价指标,揭示概念激活向量间正交化带来的二阶效应。跨概念鲁棒性衡量在对其他概念正交化后,概念检测性能的稳定性;旁侧损伤量化干预是否无意中影响了无偏见输入在下游任务上的表现。实验发现,尽管线性支持向量机具有更优的可分性和正交性,但无法实现零旁侧损伤,常落后于稀疏自编码器。在简单场景下,标准基线与基于优化的方法均无法实现完美干预。源代码即将在GitHub开源。
原文摘要 · Abstract (English)
Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existing evaluation protocols are limited to ambiguous language modeling tasks. To address this gap, we introduce SwordBench, a benchmark for steering image representations of vision models across multiple backbones and concept removal tasks. Beyond a unified benchmarking suite, we propose new evaluation notions that uncover the second-order effects of orthogonalization among concept activation vectors for pragmatic steering. Specifically, cross-concept robustness measures the stability of concept detection performance across inputs orthogonalized against alternative concepts, and collateral damage quantifies whether steering inadvertently affects model performance on a downstream task for inputs lacking the bias. We find that although a linear support vector machine exhibits superior separability and orthogonality, it fails to achieve zero collateral damage, often trailing sparse autoencoders. In simpler regimes, both standard baselines and optimization-based methods fail to achieve perfect steering. The source code will be made available soon on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。