arXiv:2608.18222cs.LGcs.CL2026-08

通过检测循环动态状态,实现测试时深度的可靠控制

Think Shallow, Solve Deep: Controlling Recurrent Dynamics for Reliable Test-Time Depth

  • 用短期动态状态判断迭代是否安全
  • 收敛态下增加深度可提效至0.34(苏格兰谜题)
  • 适合希望提升模型泛化能力的研究者

循环深度推理器通过延长测试时的迭代次数来解决更难问题,但额外迭代可能提升、保持或降低答案质量。我们发现训练后算子的有限时间动态状态(稳定、临界或漂移)能预判该行为。给出深度安全的充分条件:一旦每步位移小于解码边界,后续迭代不会改变结果。实验显示,在每难度层级800个未增强样本上训练的算法任务中,稳定态算子在增加深度时不退化,部分任务还能提升对更难未见实例的准确率(如苏格兰谜题从0.19升至0.34)。单一终端固定点目标可统一调控状态与深度行为:移除导致漂移并失去增益;加入通用递归结构后,实现进位传播任务的深度安全外推。提出四项测试时深度有效性的操作标准,用于归类失败模式,并以相同方法检验Huginn-3.5B,其属于非稳定态类别。

原文摘要 · Abstract (English)

Recurrent-depth reasoners aim to solve harder problems by iterating their update longer at test time, but additional iterations can improve, preserve, or degrade an answer. We show that a measurable property of the trained operator, its finite-time dynamical regime (estimated as settling, marginal, or drifting), indicates which of these occurs. We give a sufficient condition for depth-safety: once an operator's per-step displacement is small relative to the decoder margin, the decoded answer cannot change under further iterations. Empirically, on algorithmic tasks trained from $800$ unaugmented examples per difficulty tier, settling operators do not degrade with added depth, and on some tasks convert it into higher accuracy on harder unseen instances (Sudoku, $0.19$ to $0.34$ past the training horizon). A single terminal fixed-point objective moves the regime and the depth behavior together: removing it induces drift and removes the gains, and adding it to a generic recurrence yields depth-safe extrapolation on carry propagation. We give four operational criteria for useful test-time depth, use them to catalogue failure modes, and, as a consistency check, apply the same measurements to Huginn-3.5B, which falls in the non-settling family.

深度推理动态分析测试时优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。