让语言模型推理更稳定,深度增加也不崩溃
Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language Models

- 通过约束隐状态收敛到稳定不动点,实现可靠递归扩展
- 在数学推理任务中,深度增加时性能下降显著缓解
- 适合需要长程推理的复杂任务,如数学证明与规划
循环语言模型(LoopLMs)通过深度递归实现高效的隐空间推理,但测试时扩展行为不可靠:性能通常在某迭代深度达到峰值后随递归加深而急剧下降。通过对隐动态的分析,我们发现现有架构与策略存在稳定性与有效性之间的内在权衡。将推理视为不确定性减少过程,我们提出收敛至稳定不动点同时保持有效性的新思路。为此,提出STARS(STAbility-driven Recurrent Scaling)训练框架,通过随机循环采样实现高效的雅可比谱半径正则化,使隐状态渐近趋于稳定不动点。实验表明,在算术任务上,STARS实现了可靠的测试时扩展;在复杂数学推理任务中,显著缓解了递归深度增加带来的性能退化,同时提升了峰值表现。
原文摘要 · Abstract (English)
Looped Language Models (LoopLMs) enable efficient latent reasoning through depth recurrence, yet exhibit unreliable test-time scaling behavior: performance often peaks at a certain iteration depth and then collapses with further recurrence. Through latent dynamics analysis, we find an inherent trade-off between stability and effectiveness in existing architectures and strategies. By conceptualizing reasoning as uncertainty reduction, we propose that convergence toward stable fixed points while preserving effectiveness represents a promising way. To this end, we propose STARS (STAbility-driven Recurrent Scaling), a training framework that constrains latent states to approach asymptotically stable fixed points. This is realized via efficient Jacobian Spectral Radius Regularization with random loop sampling, enabling STARS to maximize effectiveness while ensuring rigorous stability. Experiments on arithmetic tasks show that STARS achieves reliable test-time scaling, and on complex mathematical reasoning it substantially mitigates performance degradation as recurrence depth increases while also improving peak performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。