arXiv:2509.18127cs.LGcs.AI2025-09ACL被引 5

用稀疏自编码器精细解析大模型安全特征,提升风险识别能力。

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework

  • 通过预评估指标筛选高安全解释潜力的自编码器。
  • 降低55%解释成本,实现1758个安全特征的系统化分析。
  • 开源完整工具链,助力安全研究与模型可解释性提升。

稀疏自编码器(SAEs)可通过分解模型激活来实现可解释性,但其在低频概念领域——如安全相关概念——中生成细粒度潜在特征的条件尚不明确。本文提出Safe-SAIL框架,旨在统一解析大语言模型在安全关键领域的SAE特征,深化对模型机制的理解。该框架引入预解释评估指标,高效筛选具备强安全领域可解释性的SAE,并通过分段级仿真策略将解释成本降低55%。基于此,我们训练并公开了涵盖色情、政治、暴力、恐怖四大领域共1758个安全相关特征的SAE模型集,附带可读性说明与系统评估。利用该资源,我们开展了实证分析,揭示了Safe-SAIL在风险特征识别中的有效性,并探讨了安全关键实体与概念在模型各层中的编码模式。所有模型、解释文本及工具均开源发布于配套开源工具包中。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) enable interpretability research by decomposing entangled model activations into monosemantic features. However, under what circumstances SAEs derive most fine-grained latent features for safety, a low-frequency concept domain, remains unexplored. Two key challenges exist: identifying SAEs with the greatest potential for generating safety domain-specific features, and the prohibitively high cost of detailed feature explanation. In this paper, we propose Safe-SAIL, a unified framework for interpreting SAE features in safety-critical domains to advance mechanistic understanding of large language models. Safe-SAIL introduces a pre-explanation evaluation metric to efficiently identify SAEs with strong safety domain-specific interpretability, and reduces interpretation cost by 55% through a segment-level simulation strategy. Building on Safe-SAIL, we train a comprehensive suite of SAEs with human-readable explanations and systematic evaluations for 1,758 safety-related features spanning four domains: pornography, politics, violence, and terror. Using this resource, we conduct empirical analyses and provide insights on the effectiveness of Safe-SAIL for risk feature identification and how safety-critical entities and concepts are encoded across model layers. All models, explanations, and tools are publicly released in our open-source toolkit and companion product.

可解释性安全检测稀疏编码大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。