arXiv:2601.14004cs.CL2026-01综述被引 18

提出可操作的模型可解释框架,实现定位、干预与优化闭环。

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

  • 构建'定位-引导-改进'三步流程,明确可解释对象分类
  • 在对齐性、能力与效率上实现可量化提升
  • 适合希望落地可解释性的模型优化研究者

机制可解释性(MI)已成为揭示大型语言模型(LLMs)决策过程的关键方法。然而,现有综述多将MI视为观察性科学,仅总结分析见解,缺乏系统性可操作干预框架。为此,本文提出以'定位、引导、改进'为核心的实用综述框架,基于特定可解释对象,正式分类定位(诊断)与引导(干预)方法,建立严谨的干预协议。进一步证明该框架可实现对对齐性、能力与效率的实质性提升,使MI真正成为可操作的模型优化方法。本文精选论文列表详见:https://github.com/rattlesnakey/Awesome-Actionable-MI-Survey。

原文摘要 · Abstract (English)

Mechanistic Interpretability (MI) has emerged as a vital approach to demystify the opaque decision-making of Large Language Models (LLMs). However, existing reviews primarily treat MI as an observational science, summarizing analytical insights while lacking a systematic framework for actionable intervention. To bridge this gap, we present a practical survey structured around the pipeline: "Locate, Steer, and Improve." We formally categorize Localizing (diagnosis) and Steering (intervention) methods based on specific Interpretable Objects to establish a rigorous intervention protocol. Furthermore, we demonstrate how this framework enables tangible improvements in Alignment, Capability, and Efficiency, effectively operationalizing MI as an actionable methodology for model optimization. The curated paper list of this work is available at https://github.com/rattlesnakey/Awesome-Actionable-MI-Survey.

可解释性大模型干预优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。