分层规划能提升长程控制,但需匹配低层控制器才有效
Mind the Gap: Promises and Pitfalls of Hierarchical Planning in LeWorldModel

- 冻结低层模型,高层生成隐式子目标进行规划
- 长程任务中性能提升14.7个百分点,短程则一步规划最优
- 搜索空间需贴近训练轨迹,否则会生成无效动作
我们探究时间分层能否提升LeWorldModel在长程目标条件控制中的表现。提出Hi-LeWM,冻结预训练的低层LeWM,加入高层对潜在子目标的规划。在PushT和Cube任务上测试不同目标偏移下的表现。结果表明,分层并非自动提升性能:短时程下最优配置为单步高层规划;长时程中发现高层动作空间与推理时搜索分布不匹配。使用真实未来潜变量子目标实验显示,冻结的低层控制器可准确执行对齐的中间目标,说明高层子目标生成是主要瓶颈。无约束搜索可能选择模型预测有利但实际控制效果差的宏观动作。若将搜索限制在训练轨迹编码的宏观动作附近,并合理设定子目标执行时机,可恢复有效分层机制,在中等距离提升11.3个百分点,最长PushT距离提升14.7个百分点。总体而言,时间抽象可使紧凑的冻结LeWM获益,但前提是高层搜索与低层控制器兼容。
原文摘要 · Abstract (English)
We investigate whether temporal hierarchy can improve LeWorldModel on long-horizon goal-conditioned control. We introduce Hi-LeWM, an extension that freezes the pretrained low-level LeWM and adds high-level planning over latent subgoals. We evaluate Hi-LeWM on PushT and Cube across increasing goal offsets. Hierarchy does not automatically improve performance: at short horizons, the best configuration uses a one-step high-level horizon, while longer horizons reveal a mismatch between the learned high-level action space and the inference-time search distribution. Experiments with true future latent subgoals show that the frozen low-level controller can execute well-aligned intermediate targets, indicating that high-level subgoal generation is the main bottleneck. Unconstrained search can select latent macro-actions that appear favorable under the learned model but produce poor control targets. Constraining search around macro-actions encoded from training trajectories, with appropriate subgoal execution timing, recovers useful hierarchical regimes, improving over flat LeWM by +11.3 percentage points at medium-range horizons and +14.7 percentage points at the longest PushT horizon. Overall, temporal abstraction can benefit compact frozen LeWM, but only when high-level search remains compatible with the low-level controller
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。