用分层世界模型让机器人长期任务更稳定可靠
H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model
- 分层设计:逻辑+视觉双模型协同预测状态变化
- 长序列任务误差减少,执行更稳定可靠
- 适合需要长时间规划的机器人控制场景
世界模型正成为机器人规划与控制的核心,可预测未来状态转移。现有方法多侧重视频生成或自然语言预测,难以与机器人动作对齐,且在长时序下误差累积严重。经典任务与运动规划在逻辑空间建模世界转移,具备可执行性和鲁棒性,但通常脱离视觉感知,无法同步符号与视觉状态预测。本文提出分层世界模型(H-WM),在统一框架内联合预测逻辑与视觉状态转移。该模型结合高层逻辑世界模型与低层视觉世界模型,融合符号推理的长时序鲁棒性与视觉语义的精准接地。分层输出提供稳定的中间引导,缓解误差累积,支持复杂任务序列的稳健执行。在多个视觉-语言-动作(VLA)控制策略上的实验验证了H-WM的指导效果与泛化能力。
原文摘要 · Abstract (English)
World models are becoming central to robotic planning and control as they enable prediction of future state transitions. Existing approaches often emphasize video generation or natural-language prediction, which are difficult to ground in robot actions and suffer from compounding errors over long horizons. Classic task and motion planning models world transitions in logical space, enabling robot-executable and robust long-horizon reasoning. However, they typically operate independently of visual perception, preventing synchronized symbolic and visual state prediction. We propose a Hierarchical World Model (H-WM) that jointly predicts logical and visual state transitions within a unified framework. H-WM combines a high-level logical world model with a low-level visual world model, integrating the long-horizon robustness of symbolic reasoning with visual grounding. The hierarchical outputs provide stable intermediate guidance for long-horizon tasks, mitigating error accumulation and enabling robust execution across extended task sequences. Experiments across multiple vision-language-action (VLA) control policies demonstrate the effectiveness and generality of H-WM's guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。