发现并操控了Transformer隐藏状态中的关系几何结构,提升模型推理可控性。
Relational Rank Geometry in Transformers: Detecting and Steering Hidden-State Relation Frames

- 用普鲁克符号熵检测令牌间高阶关系在隐藏空间的几何特征。
- 405B模型中86%的测试样本在预期秩处保持显著方向一致性。
- 可通过定向修复关系框架,恢复正确答案行为和残差几何结构。
Transformer隐藏状态常被解释为局部或低阶对象:神经元、稀疏特征、注意力头、残差流方向或激活块。本文研究一种互补对象:令牌元组间关系的秩索引几何结构。通过普鲁克符号熵检验r-参数关系是否在隐藏状态空间留下与基数匹配的方向签名。在Llama系列8B、70B和405B检查点中,真实关系元组在r=3,...,6时于预期秩k=r处表现出比随机打乱元组更强的方向一致性,且在多模板审计下仍具鲁棒性——所有测试的405B行均保持正向期望秩裕度,8B/70B也保留特定构造的混合单元正向裕度。进一步探究该关系几何是否可调控:在32个提示的边缘网格清洁/污染干预实验中,行/列支架与答案格式固定,仅改变是/否关系映射,将污染的隐藏状态关系框架修复至清洁或安慰剂目标。在70B和405B中,指向清洁目标的关系框架路径能恢复清洁答案行为与残差关系几何,而质心仅移位和等范数控制则无明显恢复。站点/顺序控制进一步区分标记站点重要性与有序清洁框架几何:目标清洁形状及跨提示清洁形状可在标记接口恢复行为与残差几何,而污染源转移、同站点置换/反射、错误站点清洁差分、质心仅移动与等范数噪声均无效或远低于清洁框架路径。结果建立了一条从关系探测到关系框架干预的可控桥梁:关系秩几何可在Transformer隐藏状态中被检测、靶向并行为验证。
原文摘要 · Abstract (English)
Transformer hidden states are often interpreted through local or low-order objects: neurons, sparse features, attention heads, residual-stream directions, or activation patches. This paper studies a complementary object: the rank-indexed geometry of relations among token tuples. I use Plucker sign entropy to test whether r-argument relations leave arity-matched orientation signatures in hidden-state space. Across Llama-family 8B, 70B, and 405B checkpoints, true relation tuples show stronger orientation-sign consistency at the expected rank k=r for r=3,...,6 than scrambled tuples under matched random-control audits. Multi-template audits show that the effects survive surface variation, with all tested 405B rows retaining positive expected-rank margins and 8B/70B retaining positive rows with constructor-specific mixed cells. I then ask whether the same relation geometry can be steered. In an edge-grid clean/corrupt intervention assay over 32 prompts, the row/column scaffold and answer format stay fixed while the YES/NO relation map changes, and the corrupt hidden-state relation frame is patched toward clean or placebo targets. In 70B and 405B, clean-targeted relation-frame paths recover clean-answer behavior and residual relation geometry, while centroid-only and equal-norm controls show negligible recovery. Site/order controls further separate marker-site importance from ordered clean-frame geometry: target clean shape and cross-prompt clean shape recover behavior and residual geometry at the marker interface, whereas corrupt-donor transfer, same-site permutation/reflection, wrong-site clean deltas, centroid-only motion, and equal-norm noise fail or remain far below clean-frame paths. The result is a controlled bridge from relation probing to relation-frame intervention: relation rank geometry can be detected, targeted, and behaviorally validated in transformer hidden states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。