用分层世界模型让机器人实时执行复杂时序指令,兼顾安全与进度。
hint$^2$: Hierarchical World Models for Inference-Time Temporal Logic Guidance

- 分两级建模:高层预测任务进展,低层确保局部动作安全。
- 在CALVIN数据集上超越现有方法,成功完成含强约束的复杂指令。
- 适合需要高可靠性的机器人推理阶段控制,如工业操作场景。
机器人学习的核心目标之一是实现在运行时执行丰富的动态指令。尽管大规模语言条件策略已取得显著进展,但在处理时序结构和安全约束方面仍存在困难。线性时序逻辑(LTL)能有效表达复杂的非马尔可夫指令,但如何引导学习到的操纵策略满足LTL规范仍具挑战性,因现代策略生成短时域动作片段并闭环重规划,而绝大多数LTL规范需在长时域轨迹上评估。本文提出hint²,一种基于分层世界模型的推理时时间逻辑引导方法。核心思想是利用不同抽象层级的世界模型生成双重引导目标:高层模型预测动作引发的任务相关原子命题转移,以推进LTL自动机状态;低层动力学模型预测即时状态演化,实现精确局部安全引导。实验表明,hint²克服了现有LTL引导扩散方法的局限,在CALVIN基准上优于已有推理时调制方法,并更优雅地完成包含复杂活度与安全约束的指令。最后,我们验证了hint²可在真实UR5e机械臂上执行复杂指令。
原文摘要 · Abstract (English)
A central goal of robot learning is to enable robots to execute rich instructions specified at runtime. Large-scale language-conditioned policies have made substantial progress toward this goal, yet still struggle with temporal structure and safety constraints. Linear Temporal Logic (LTL) provides a powerful language to express complex, non-Markovian instructions. However, guiding learned manipulation policies toward LTL satisfaction remains challenging because modern policies generate short-horizon action chunks and replan in closed loop, while almost all LTL specifications are evaluated over long-horizon trajectories. In this paper, we introduce hint$^2$, a method for guiding short-horizon policies toward satisfying complex LTL specifications at inference time using hierarchical world models. Our key idea is to derive two separate guidance objectives using each world model's abstraction level. A high-level model predicts future action-induced transitions in task-relevant atomic propositions to guide progress through the LTL automaton, while a low-level dynamics model predicts immediate state evolution for accurate local safety guidance. Our results show that hint$^2$ overcomes the limitations of current LTL-guided diffusion methods, outperforms existing inference-time steering methods in CALVIN, and successfully completes instructions with complex liveness and safety constraints more elegantly than language-conditioned alternatives. Finally, we demonstrate that hint$^2$ can handle complex instructions on a real UR5e manipulator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。