arXiv:2511.20820cs.CL2025-11Conference of the …被引 10

用智能体框架主动解释语言模型的稀疏编码特征

SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models

  • 构建智能体系统,通过实验验证和迭代优化解释
  • 生成与预测准确率显著优于现有方法
  • 适合研究模型可解释性与内部机制的学者

大语言模型(LLMs)虽取得显著进展,但其内部机制仍不透明,制约其安全可靠部署。稀疏自编码器(SAEs)能将模型表示分解为更易理解的特征,但解释这些特征仍具挑战。本文提出SAGE(SAE AGentic Explainer),一种基于智能体的框架,将特征解释从被动单次生成转变为主动、以解释为导向的过程。SAGE通过系统化提出多个解释、设计针对性实验验证,并基于实测激活反馈迭代优化解释。在多种语言模型的SAE特征上进行实验表明,SAGE生成的解释在生成与预测准确性上显著优于当前最优基线。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved remarkable progress, yet their internal mechanisms remain largely opaque, posing a significant challenge to their safe and reliable deployment. Sparse autoencoders (SAEs) have emerged as a promising tool for decomposing LLM representations into more interpretable features, but explaining the features captured by SAEs remains a challenging task. In this work, we propose SAGE (SAE AGentic Explainer), an agent-based framework that recasts feature interpretation from a passive, single-pass generation task into an active, explanation-driven process. SAGE implements a rigorous methodology by systematically formulating multiple explanations for each feature, designing targeted experiments to test them, and iteratively refining explanations based on empirical activation feedback. Experiments on features from SAEs of diverse language models demonstrate that SAGE produces explanations with significantly higher generative and predictive accuracy compared to state-of-the-art baselines.an agent-based framework that recasts feature interpretation from a passive, single-pass generation task into an active, explanationdriven process. SAGE implements a rigorous methodology by systematically formulating multiple explanations for each feature, designing targeted experiments to test them, and iteratively refining explanations based on empirical activation feedback. Experiments on features from SAEs of diverse language models demonstrate that SAGE produces explanations with significantly higher generative and predictive accuracy compared to state-of-the-art baselines.

可解释性稀疏编码智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。