arXiv:2506.19382cs.CL2025-06NeurIPS被引 12

提出新指标与方法,让大模型内部特征更清晰可控。

Measuring and Guiding Monosemanticity

  • 用新指标FMS量化特征单一性,精准评估模型内部表示
  • 训练时引入标签引导,使目标概念在隐空间更独立分明
  • 在毒性检测等任务中提升可解释性,适合研究模型机制者

当前方法在定位和操控大语言模型(LLMs)内部特征表示时面临根本性挑战。稀疏自编码器(SAEs)虽为大规模特征提取提供了新方向,但仍受限于特征隔离不完全和不可靠的单义性。为此,我们引入特征单义性评分(FMS),系统量化隐表示中的特征单义性。基于此,我们提出有指导的稀疏自编码器(G-SAE),在训练中以标注概念为条件调节隐表示。实验表明,隐空间中目标概念的可靠定位与解耦显著提升了可解释性、行为检测能力与控制精度。在毒性检测、写作风格识别和隐私属性识别任务中,G-SAE不仅增强单义性,还实现了更有效、更精细的调控,且性能下降更小。研究结果为推进机制可解释性与控制提供了可操作的指导。

原文摘要 · Abstract (English)

There is growing interest in leveraging mechanistic interpretability and controllability to better understand and influence the internal dynamics of large language models (LLMs). However, current methods face fundamental challenges in reliably localizing and manipulating feature representations. Sparse Autoencoders (SAEs) have recently emerged as a promising direction for feature extraction at scale, yet they, too, are limited by incomplete feature isolation and unreliable monosemanticity. To systematically quantify these limitations, we introduce Feature Monosemanticity Score (FMS), a novel metric to quantify feature monosemanticity in latent representation. Building on these insights, we propose Guided Sparse Autoencoders (G-SAE), a method that conditions latent representations on labeled concepts during training. We demonstrate that reliable localization and disentanglement of target concepts within the latent space improve interpretability, detection of behavior, and control. Specifically, our evaluations on toxicity detection, writing style identification, and privacy attribute recognition show that G-SAE not only enhances monosemanticity but also enables more effective and fine-grained steering with less quality degradation. Our findings provide actionable guidelines for measuring and advancing mechanistic interpretability and control of LLMs.

可解释性自编码器特征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。