arXiv:2605.28649cs.LGcs.CL2026-05

用SAE诊断模型层,而非直接修改,提升数学推理能力

Interpretability-Guided Layer Selection over Subspace Projection: SAEs as Stethoscopes, Not Scalpels, for Raw Task Vector Model Editing

论文配图:Interpretability-Guided Layer Selection over Subspace Projection: SAEs as Stethoscopes, Not Scalpels, for Raw Task Vector Model Editing
图 1 · 摘自论文原文
  • 将任务向量注入SAE识别的特定层,不进行特征过滤
  • 数论准确率从29.6%提升至39.4%,5/7科目显著改善
  • 适合追求可解释性编辑的模型优化研究者

大型语言模型需高效编辑以增强领域能力,避免全量微调的开销与灾难性遗忘。稀疏自编码器(SAEs)被用于定位干预点,但本研究在Gemma-3-4B-IT上评估其指导的数学推理编辑流程时发现:将任务向量投影到SAE特征子空间会丢弃约97%的修改能量,导致七类数学任务均无统计显著提升。根源在于激活空间方向与权重空间任务向量存在几何错位。我们提出新视角:将SAE视为‘听诊器’而非‘手术刀’,仅用于层级诊断。通过将未过滤的任务向量注入由SAE特异性分数识别的层,数论准确率从29.6%升至39.4%(z=+3.41, p=0.0007),5/7科目显著提升,无一显著下降。方法完全确定、无需额外推理成本,提供可解释编辑的理论框架。

原文摘要 · Abstract (English)

LLMs increasingly require surgical model editing to enhance domain-specific capabilities without incurring the computational cost or catastrophic forgetting associated with full fine-tuning. Sparse Autoencoders (SAEs) have emerged as a promising tool in this setting, in principle allowing for feature-level identification of where to intervene. In this work, we rigorously evaluate an SAE-guided editing pipeline for mathematical reasoning on Gemma-3-4B-IT and uncover a fundamental failure mode: the intuitively appealing approach of projecting task vectors onto SAE feature subspaces acts as an information bottleneck that discards approximately 97% of the modification energy, yielding no statistically significant improvements across seven math subjects. We show that this failure stems from a geometric misalignment between activation-space SAE directions and weight-space task vectors. We then propose a shift in perspective: SAE as a Stethoscope, Not a Scalpel, where SAEs are used for layer-level diagnosis rather than intervention-level filtering. By injecting unfiltered raw task vectors only into layers identified by an SAE-derived specificity score, we improve Number Theory accuracy from 29.6% to 39.4% (z=+3.41, p=0.0007) on the Minerva Math benchmark; 5 of 7 math subjects significantly improved and none significantly degraded. Our method is fully deterministic, requires no additional inference cost, and provides a principled framework for interpretability-guided model editing.

模型编辑可解释性稀疏自编码器数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。