用语言预测中间目标,实现长程精准规划
Latent Goal Prediction from Language for Model-Based Planning

- 从语言指令中预测潜在空间的中间目标序列
- 在多环境中实现无退化长程规划,优于以往方法
- 适合需要语言引导、高采样率的模型规划任务
基于世界模型的规划受限于累积预测误差以及可优化目标定义困难。视觉目标提供精确局部梯度但缺乏远距离指导,语言灵活却受跨模态噪声或依赖大型生成模型影响,不适用于模型规划所需的高采样特性。为此,我们提出从语言预测潜在目标(LAGO)框架,能在同一潜在空间内预测语言指令对应的中间目标状态序列及动作条件滚动路径。不同于优化单一全局目标,LAGO动态将指令分解为显式预测的局部可处理潜在子目标。通过在线更新子目标并使用软最小轨迹成本进行规划,使智能体能沿长时程保持一致的潜在轨迹。在多个环境与规划跨度上的评估表明,LAGO避免了先前方法的性能急剧下降。仅凭语言即可实现稳健且精确的长程规划,融合了视觉目标的精度与文本引导控制的灵活性。
原文摘要 · Abstract (English)
Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals. Visual targets provide precise local gradients but poor distant guidance, while language is flexible yet limited by noisy cross-modal alignment or dependence on large generative models unsuited for the high-sampling nature of model-based planning. To address these challenges, we introduce Latent Goal Prediction from Language (LAGO), a framework that predicts both sequences of intermediate goal states from language instructions and action-conditioned rollouts, all within the same latent space. Rather than optimizing toward a single global objective, LAGO dynamically decomposes instructions into explicitly predicted, locally tractable latent subgoals. By updating these subgoals online and using a soft minimum trajectory cost during planning, LAGO enables an agent to follow coherent latent trajectories over long horizons. Evaluation across multiple environments planning horizons shows that LAGO avoids the sharp degradation of prior methods. By achieving robust and precise long-horizon planning purely from language, LAGO bridges the precision of visual goals with the flexibility of text-guided control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。