arXiv:2606.03741cs.AI2026-06中稿 · the Workshop on Co…被引 1

在隐式推理中,合理保持子目标能显著提升长程规划效果。

When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning

  • 用高层管理器周期性发出持续P步的子目标,引导底层推理
  • 子目标持续期P=3时模型损失最低(1.544),优于频繁或长期保持
  • 适配的对齐权重λ≈0.05,过强会破坏已学习的方向结构

长程推理需在稳定性与适应性间权衡:重规划过于频繁则计算无法形成多步结构,保持过久又会导致计划过时。本文研究隐式推理中的这一权衡,其中多步计算发生在隐藏状态内部而非外部标记序列。扩展层次推理模型(HRM)为类封建式管理者-工作者架构:慢速高层模块定期发出归一化方向子目标,持续作用于底层P步,通过隐式余弦对齐损失引导工作者隐藏状态更新。在ARC和ConceptARC数据集上,发现子目标持久性(非仅注入)是核心调控参数:中等持续期P∈[3,6] consistently表现最优,尤其在P=3时损失最低(1.544,对比P=1时1.674,基线1.640;5次种子均值1.595±0.045)。对齐权重λ存在互补窄优解(约0.05)。控制消融实验表明,当对齐信号超过最优值时,干扰来源是已学习的方向结构,而非架构容量或辅助损失本身。结果揭示隐式推理系统中组合规划的设计原则:中等时域意图需在足够计算步骤内保持一致,以促成组合结构形成。

原文摘要 · Abstract (English)

Long-horizon reasoning requires a system to commit to medium-horizon intent without becoming rigid: re-plan too often and computation never coheres into multi-step structure; commit too long and the plan goes stale. We study this stability-adaptivity tradeoff in the latent reasoning setting, where multi-step computation occurs inside hidden state rather than externalized token traces. We extend the Hierarchical Reasoning Model (HRM) with a feudal-style manager-worker interface: a slow high-level module periodically emits a normalized directional subgoal that persists for P low-level steps, biasing the worker's hidden-state updates and supplying an intrinsic cosine alignment loss. On ARC and ConceptARC, we find that subgoal persistence -- not subgoal injection alone -- is the central knob: moderate periods P in [3, 6] consistently outperform both very frequent (P=1) and very long horizons, with a clear minimum LM loss at P=3 (1.544 vs. 1.674 at P=1, 1.640 baseline; replicated over 5 seeds at mean 1.595, std 0.045). The intrinsic alignment weight lambda shows a complementary narrow optimum (lambda approximately 0.05). A controlled ablation at past-sweet-spot lambda isolates learned directional structure -- not architectural capacity or auxiliary loss alone -- as the source of interference when the alignment signal exceeds its optimum. Together these findings implicate a design principle for compositional planning in latent reasoning systems: medium-horizon intent must be coherent across enough computational steps for compositional structure to form.

隐式推理层次规划子目标长程决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。