语言模型能自我评估计算质量,外部干预可提升结果。
Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary
- 通过循环变压器结构检测计算质量,用隐藏状态预测成功概率。
- 在GSM8K任务中,预测准确率提升0.066,跨任务验证有效。
- 适合研究模型可解释性与可控推理的学者参考。
我们测试了一个冻结的2.6B参数循环Transformer模型Ouro-RLTT是否能读取自身计算质量,并判断外部干预能否带来更好结果。在GSM8K上,预答案探针虽排除答案区域和真值,仍能预测成功:隐藏状态结合长度/对数概率特征的AUROC达0.797,优于仅用表面特征的0.731(提升+0.066;任务聚类95%置信区间[+0.021,+0.112],170个任务)。在Horizon Logic任务中,前瞻扩展的不重叠任务研究显示提升+0.111(置信区间[+0.056,+0.169]),并在新数据集上独立复现(+0.095),且对对抗性错误捷径具有鲁棒性。循环机制使候选质量可读性逐步提前至更浅层物理深度;该趋势在Ouro系列中重复出现,且在异族模型Huginn中定性一致,尽管转移几何不同。基于隐藏状态的评分在四个封闭选择预测任务中优于仅依赖捷径的评分,终端选择在所有候选均合理时仍优于随机(27/32正确,预期64.8%;p=0.0086)。生成控制无效:方向引导为负,分支筛选有界,精确计算循环分配与最小LoRA方向绑定均未见收益。所有测试在192槽循环缓存上通过比特精确分支/进位/剪枝机制运行,包括后缀重计算拼接,最多节省88%每分支层计算量。我们称此为可决策使用但不可生成控制的特性为操作性前内省。所有关键指标采用源项不重叠划分与反对称成对评估。
原文摘要 · Abstract (English)
Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes? We test both questions in a frozen 2.6B looped transformer, Ouro-RLTT. On GSM8K, a strict pre-answer probe excludes the answer region and gold value yet predicts success: hidden states plus length/log-probability features reach AUROC 0.797 versus 0.731 for those surface features alone (increment +0.066; task-clustered 95% CI [+0.021,+0.112]; 170 tasks). On Horizon Logic, a prospectively extended task-disjoint study gives an increment of +0.111 (CI [+0.056,+0.169]), independently replicated on the new cohort (+0.095) and robust to an adversarial malformed-sibling shortcut. Recurrence also moves candidate-quality readability to progressively earlier physical depth; the trend replicates across the Ouro family and qualitatively in out-of-family Huginn, although their transfer geometry differs. The readout converts into validated decision-level gains. Hidden-state-based scores improve risk-coverage over shortcut-only scores in four sealed selective-prediction arms, and terminal selection beats matched random even when every candidate is well formed (27/32 correct selections versus 64.8% expected; p = 0.0086). Generative control does not convert: directional steering is negative, a branch screen is bounded, and exact-compute loop allocation and minimal LoRA direction-binding detect no gain. These tests run through bit-exact branch/carry/prune machinery over Ouro's 192-slot recurrent cache, including a suffix-recompute splice saving up to 88% of per-branch layer passes. We call this decision-usable but not generatively controllable property operational proto-introspection. All load-bearing values use source-item-disjoint splits and antisymmetrized pairwise evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。