arXiv:2608.13626cs.AIcs.CL2026-08

发现状态信号可独立于全局闭包存在,验证了动作映射的局部可解性。

A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure

论文配图:A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure
图 1 · 摘自论文原文
  • 通过几何格验证,无源动作映射仍能通过单步、复合等门控测试
  • 23/30最强单元在闭包门上翻转,表明校准非普遍成立
  • 适用于研究模型内部状态与动作映射关系的读者

隐藏状态信号可在不支持可复用动作映射的情况下仍可解码或因果使用。我们测试了在无源情况下,动作映射是否能达到自然后动作激活和组合。在已知仿射S_5载体上构建证据格,所有持源折叠均通过单步、复合、逆元、解码及交换性门控。结构曲率与持域共轭导致误差单调上升,但仅23/30最强单元翻转闭包门,表明校准非普遍。在后训练Qwen/Qwen3-4B中,冻结末标记h28仿射映射的平均持实体误差为0.519,而域内交叉拟合为0.398。七次随机实体划分与映射几何不支持纯实体特异性解释。早期层h4/h16拟合单步转移更优,但h4冲突态解码弱,词汇控制仍未解决。三组从同一冻结检查点重构的干预数据集显示,因果效应仅出现在h28/h36。结果感知重拟合使h28单步误差降至0.474(加权为0.469),但无重拟合通过复合门。学习有限世界同样保持相对代数信号或共享图表,但无持源仿射闭包。在测试载体中,状态可用性、因果使用、局部几何与可复用闭包可分离。结论限于一个预训练模型、采样末标记层、两个有限世界及测试仿射或诊断函数类。

原文摘要 · Abstract (English)

A hidden state signal can be decodable or causally usable without supporting a reusable action map. We test whether action maps fitted without a source reach its natural post-action activation and compose. We organize the tests as an evidence lattice and validate the geometric branch on a known affine S_5 carrier: all held-source folds pass one-step, composition, inverse, decoding, and commutativity gates. Structured curvature and held-domain conjugacy raise error monotonically, but only 23/30 strongest cells flip a closure gate, bounding rather than universalizing calibration. In post-trained Qwen/Qwen3-4B, frozen final-token h28 affine maps have mean held-entity error .519, versus .398 for within-test-domain cross-fit. Seven randomized entity splits and map geometry do not support a purely entity-specific account. Earlier h4/h16 layers fit one-step transitions better, but h4 conflict-state decoding is weak and lexical controls remain unresolved. Three matched intervention datasets regenerated from one frozen checkpoint show causal effects only at h28/h36. Outcome-aware refitting improves h28 one-step error to .474 (.469 with weighting), yet no refit passes composition. Learned finite worlds likewise preserve relative algebraic signals or shared charts without held-source affine closure. Within the tested carriers, state availability, causal use, local geometry, and reusable closure are separable. The result is limited to one pretrained model, sampled final-token layers, two finite worlds, and the tested affine or diagnostic function classes.

状态信号动作映射仿射闭包模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。