arXiv:2608.11822cs.CLcs.LG2026-08

发现大模型中存在被抑制的推理结构,可定位但难释放。

Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release

  • 通过干预中间层证据通道实现行为恢复,定位成功。
  • 检测器在分布外失效,导致门控机制完全倒置。
  • 线性释放效果受限,增益不足0.08阈值,且无法自适应提升。

大量研究指出语言模型包含任务相关潜在结构却未使用。这些结构能否被定位并转化为行为,尚缺乏端到端验证。本文对一个2570万参数的Transformer模型进行全预注册压力测试,该模型训练于因果证据判别任务,此前已知存在潜在因果结构被抑制的现象。所有阈值、判断模板与决策树分支均在数据生成前哈希归档。结果:(i) 定位成功——在中间层观测-证据通道施加干预,可在原本被抑制的世界中恢复目标行为(配对释放优势分别为0.563和0.854,97.5%置信区间不包含零;最佳位置释放率为0.889);(ii) 门控机制在分布外失效——校准于分布外世界的检测器在分布内生成中触发率达6.9–7.3%,但在2400个真正需要释放的生成中未触发,导致整个门控流程退化为基线模型;(iii) 线性释放被限制——移除门控并直接注入每实例线性方向,呈现单调剂量反应,但平台远低于预注册释放阈值(截距从0.382降至0.264,对比阈值≤0.08);实例自适应增益小于±0.03。失败双重定位:检测器分布外倒置,且该位置与分辨率下的线性释放方向族整体远离充分性。两项失败可分离,但不否定定位有效性。所有数值均可追溯至公开审计链中的哈希记录。

原文摘要 · Abstract (English)

A growing body of work reports that language models represent task-relevant latent structure that they fail to use. Whether such structure, once located, can be converted into behavior is a separate question that is rarely tested end to end. We submit the complete pipeline -- detect, localize, and release -- to a fully preregistered stress test on a 25.7M transformer trained on causal-evidence discrimination, where a known suppression phenomenon (latent causal structure present but behaviorally unused) has previously been documented. Every threshold, claim template, and decision-tree branch was hashed and archived before any corresponding data existed. Three findings. (i) Localization succeeds: interventions at observation-evidence channels of mid layers restore target behavior on otherwise-suppressed worlds (paired release advantages $0.563$ and $0.854$, 97.5% CIs excluding zero; best-site release rate $0.889$). (ii) Gating fails out of distribution: a detector calibrated to trigger on zero out-of-distribution calibration worlds triggers on 6.9-7.3% of held-out in-distribution generations and on zero of the 2,400 held-out generations that actually need it -- a complete inversion that silently reduces the gated pipeline to its base model. (iii) Linear release is capped: removing the gate and injecting a per-instance linear direction unconditionally yields a monotone dose-response that plateaus far below the preregistered release margin (intercept $0.382 \to 0.311 \to 0.264$ vs. threshold $\le 0.08$); per-instance adaptivity adds less than $\pm 0.03$. The failure is doubly located: the detector is OOD-inverted, and the entire family of linear release directions at this site and resolution is bounded away from sufficiency. The two failures are dissociable, and neither overturns localization. Every number traces to a hashed artifact in the released audit chain.

大模型可解释性门控机制行为释放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。