提出新方法诊断语言模型修复是否真有效,避免假象干扰。
Evaluating the Semantic Specificity of Representation Steering in Language Models

- 用模型本就擅长的规则族测试干预效果,识别真实修复
- 发现干预使正确率从99.6%跌至40.4%,实为强加错误标签
- 提供四重验证,适合关注模型可解释性与可靠性研究者
局部表征操纵(LRS)被广泛用于修正大语言模型的推理缺陷。然而,标准评估易受表面标签覆盖误导,造成修复假象。本文提出跨规则迁移(CRT)诊断框架,通过在模型原本擅长的规则族上评估干预效果来审计表征操作。针对普遍存在的矛盾盲视问题,对深层层LRS的评估显示,该干预仅引入全局标签偏差:将操纵向量应用于模型本已正确处理的规则(基线正确率99.6%),导致性能降至40.4%,强制产生错误矛盾判断。我们通过四项互补控制实验(直接逻辑偏置等价、控制向量标签翻转、跨模型嫁接、早期层操纵检查)支持此诊断,建立了一套严谨方法,以区分真正推理修复与表面标签覆盖。
原文摘要 · Abstract (English)
Localized Representation Steering (LRS) is widely used to correct reasoning pathologies in large language models. However, standard benchmark evaluations can easily be fooled by superficial label overrides, creating a false impression of reasoning circuit repairs. In this work, we propose Cross-Rule Transfer (CRT), a diagnostic framework that audits representational interventions by evaluating them on rule families where the model is natively competent. Evaluating late-layer LRS for a widespread logical failure, contradiction blindness, reveals that the intervention merely injects a global label bias: applying the steering vector to rules the model already handles correctly (99.6% baseline) degrades performance to 40.4% by forcing false contradiction predictions. We support this diagnosis with four complementary controls (direct logit bias equivalence, control vector label-flipping, cross-model grafting, and early-layer steering checks), providing a rigorous methodology to distinguish genuine reasoning repairs from superficial label overrides.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。