arXiv:2512.02004cs.LGcs.CL2025-12被引 4

让大模型的隐藏特征对齐人类概念,实现精准控制与解释。

AlignSAE: Concept-Aligned Sparse Autoencoders

  • 先无监督训练再用概念监督微调,将特定概念绑定到独立编码槽。
  • 单个概念槽可实现稳定的概念替换,支持多跳推理和泛化机制分析。
  • 适合需要可解释性、可控性的大模型研究者与应用开发者。

大型语言模型(LLMs)将事实知识编码在难以观测或调控的隐层参数空间中。稀疏自编码器(SAEs)虽能将隐藏激活分解为更细粒度、可解释的特征,但常无法可靠地将这些特征与人类定义的概念对齐,导致特征表示纠缠且分散。为此,我们提出AlignSAE,通过“预训练-后训练”教学流程,将SAE特征与预定义本体对齐。初始无监督训练后,采用监督后训练将特定概念绑定至专用潜变量槽,同时保留其余容量用于通用重构。这种分离机制构建了一个可解释的接口,使特定概念可被独立检查与控制,不受无关特征干扰。实证结果表明,AlignSAE支持精确因果干预,如可靠的“概念替换”,仅需操作单一语义对齐的潜变量槽,并进一步支持多跳推理及对类似‘领悟’现象的机制探测。

原文摘要 · Abstract (English)

Large Language Models (LLMs) encode factual knowledge within hidden parametric spaces that are difficult to inspect or control. While Sparse Autoencoders (SAEs) can decompose hidden activations into more fine-grained, interpretable features, they often struggle to reliably align these features with human-defined concepts, resulting in entangled and distributed feature representations. To address this, we introduce AlignSAE, a method that aligns SAE features with a predefined ontology through a "pre-train, then post-train" curriculum. After an initial unsupervised training phase, we apply supervised post-training to bind specific concepts to dedicated latent slots while preserving the remaining capacity for general reconstruction. This separation creates an interpretable interface where specific concepts can be inspected and controlled without interference from unrelated features. Empirical results demonstrate that AlignSAE enables precise causal interventions, such as reliable "concept swaps", by targeting single, semantically aligned slots, and further supports multi-hop reasoning and a mechanistic probe of grokking-like generalization dynamics.

可解释性稀疏编码大模型控制概念对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。