发现稀疏自编码器的概念表示易受微小干扰影响,可能误导模型监控。
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
- 将鲁棒性建模为输入空间优化问题,设计对抗扰动测试框架
- 微小扰动即可改变概念解释,而基础模型激活基本不变
- 提醒研究者:概念表示需额外去噪,否则不适用于模型监管
稀疏自编码器(SAEs)常用于将大语言模型(LLMs)的内部激活映射为人类可理解的概念表示。现有评估多关注重建与稀疏性的权衡、人工(自动)可解释性及特征解耦,却忽略了概念表示对输入扰动的鲁棒性这一关键问题。我们主张,鲁棒性应是概念表示的基本要求,反映概念标注的可靠性。为此,我们将鲁棒性量化建模为输入空间优化问题,构建了一个包含真实场景的综合性评估框架,通过生成对抗性扰动来操纵SAE表示。实验表明,在大多数情况下,微小的对抗性输入扰动可有效改变基于概念的解释,而基础LLM的激活几乎不受影响。整体结果表明,SAE概念表示脆弱,若无进一步去噪或后处理,可能不适用于模型监控与监督等应用。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) are commonly used to interpret the internal activations of large language models (LLMs) by mapping them to human-interpretable concept representations. While existing evaluations of SAEs focus on metrics such as the reconstruction-sparsity tradeoff, human (auto-)interpretability, and feature disentanglement, they overlook a critical aspect: the robustness of concept representations to input perturbations. We argue that robustness must be a fundamental consideration for concept representations, reflecting the fidelity of concept labeling. To this end, we formulate robustness quantification as input-space optimization problems and develop a comprehensive evaluation framework featuring realistic scenarios in which adversarial perturbations are crafted to manipulate SAE representations. Empirically, we find that tiny adversarial input perturbations can effectively manipulate concept-based interpretations in most scenarios without notably affecting the base LLM's activations. Overall, our results suggest that SAE concept representations are fragile and without further denoising or postprocessing they might be ill-suited for applications in model monitoring and oversight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。