用稀疏自编码器精准操控大模型的高阶语义特征。
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
- 通过对比语义对立构建检索管道,从稀疏激活空间中提取单一语义功能特征。
- 在五大性格特质上实现双向精确控制,稳定性与性能优于现有方法。
- 发现功能忠实性现象:干预单一特征可引发多维度语义协同变化,适合可控生成研究者。
机制可解释性(MI)近年实现了对大语言模型(LLM)内部特征的识别与干预,但如何将这些内部特征与语言生成中的复杂行为级语义属性可靠关联仍是一大挑战。本文提出一种基于稀疏自编码器的框架,用于检索和操控与高层语言行为相关的可解释内部特征。方法采用基于受控语义对立的对比特征检索流程,结合统计激活分析与生成验证,从稀疏激活空间中提炼出单义性的功能性特征。以五大性格特质为案例,证明该方法可在保持优异稳定性和性能的前提下,实现模型行为的精准双向调控,优于现有的对比激活添加(CAA)等方法。进一步发现一种称作‘功能忠实性’的实证效应:对特定内部特征进行干预,会引发多个语言维度的协调且可预测的语义迁移,与目标语义属性一致。结果表明,LLMs 内部蕴含高度整合的高阶概念表征,并为复杂人工智能行为的机制化调控提供了新路径。
原文摘要 · Abstract (English)
Recent work in Mechanistic Interpretability (MI) has enabled the identification and intervention of internal features in Large Language Models (LLMs). However, a persistent challenge lies in linking such internal features to the reliable control of complex, behavior-level semantic attributes in language generation. In this paper, we propose a Sparse Autoencoder-based framework for retrieving and steering semantically interpretable internal features associated with high-level linguistic behaviors. Our method employs a contrastive feature retrieval pipeline based on controlled semantic oppositions, combing statistical activation analysis and generation-based validation to distill monosemantic functional features from sparse activation spaces. Using the Big Five personality traits as a case study, we demonstrate that our method enables precise, bidirectional steering of model behavior while maintaining superior stability and performance compared to existing activation steering methods like Contrastive Activation Addition (CAA). We further identify an empirical effect, which we term Functional Faithfulness, whereby intervening on a specific internal feature induces coherent and predictable shifts across multiple linguistic dimensions aligned with the target semantic attribute. Our findings suggest that LLMs internalize deeply integrated representations of high-order concepts, and provide a novel, robust mechanistic path for the regulation of complex AI behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。