发现循环Transformer内部状态中关系性偏好编码更有效,但原结果因数据泄漏被修正。
Relational Preference Encoding in Looped Transformer Internal States

- 通过关系性对比而非逐点判断,可更准确解码人类偏好
- 修正后配对准确率仅56.5%,低于原报告的84.5%;点判准确率54.2%
- 揭示数据泄漏与排序偏差双重陷阱,需联合检测
我们研究循环Transformer如何编码人类偏好,在冻结的Ouro-2.6B模型迭代状态上训练轻量级评估头,基于Anthropic HH-RLHF数据集。原报告中的三项核心结果因独立评估错误被高估:95.2%的成对评估准确率实为顺序优先偏差所致;全集8,552对测试中反称化准确率为63.9%。成对探测器原报84.5%,点判探测器21.75%(低于随机),均因源项泄露导致;修正后配对与点判准确率分别为56.5%和54.2%。核心结论仍成立:关系性解码优于点判(+2.3分,95%置信区间[+1.3, +3.3]),反称化评估器仍优于线性探测器,但无修正读出能超越端到端奖励模型。方法论发现(恒定输出退化、翻转测试、交换协议度量压缩)成立,唯一修正:反称化准确率而非反称相关性才可验证关系区分能力。两种错误相互隐蔽——拆分审计无法察觉顺序先验,反称化无法察觉泄露——故二者缺一不可。完整审计见后续工作(Kirin, 2026,待发表)。
原文摘要 · Abstract (English)
We investigate how looped transformers encode human preference, training lightweight evaluator heads on frozen Ouro-2.6B loop-iteration states on Anthropic HH-RLHF. v2: an erratum is prepended; the original manuscript is unchanged. A post-publication audit found the three headline results inflated by two independent evaluation errors. The 95.2% pairwise evaluator accuracy is a canonical-ordering artifact: the data were correctly split, but the evaluator learned to prefer the first-presented argument; its strict antisymmetrized accuracy on the full 8,552-pair test set is 63.9%. The 84.5% pairwise probe and the below-chance 21.75% pointwise probe were source-item leaks (orientation rows and pair partners crossing the train/test split); corrected pair-disjoint values are 56.5% and 54.2% -- above chance, so the "inverted polarity" finding is withdrawn. The central finding survives at much smaller magnitude: preference is decoded more accurately relationally than pointwise (paired +2.3 points, 95% CI [+1.3, +3.3]), and the antisymmetrized evaluator still beats the linear probe, but no corrected readout rivals end-to-end reward models. The methodological findings stand (constant-output degeneracy, flip test, swap-protocol metric deflation), with one correction: antisymmetrized accuracy, not antisymmetry correlation, certifies relational discrimination. The two errors are mutually invisible -- a split audit cannot see an ordering prior, antisymmetrization cannot see a leak -- so both checks are required. Full audit in the follow-up work (Kirin, 2026, in preparation).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。