用任务谓词验证未来,让机器人规划更可靠。
EV-WM: Event-Verified World Models for Long-Horizon Robotic Manipulation

- 在预训练特征空间中滚动预测未来,解析为结构化事件状态
- 通过任务进展、语义一致性和物理可行性评分筛选可行路径
- 适合长时程、接触敏感的复杂操作任务,提升规划可解释性
预训练特征世界模型为机器人想象提供了基础,但仅靠视觉或潜在空间预测无法判断想象的未来是否满足任务相关条件。长时程操作需要关系性、谓词级且物理可信的进展信号:物体是否移动、抽屉或接触状态是否改变、放置谓词是否满足、候选未来是否足够可靠以执行。我们提出EV-WM,一种基于谓词的验证框架,用于世界模型规划。EV-WM在预训练视觉特征空间中滚动生成候选未来,将其解码为结构化事件状态,并利用任务进展、语义一致性、物理可行性及不确定性项进行评分。该验证器指导基于采样的规划,筛选候选动作,在接触敏感的LIBERO酒架场景中从PPO生成的提议中选出最优方案。在导航、可变形物体、墙约束及语言描述的操作任务中,实验表明,基于谓词的验证能显著提升特征空间世界模型规划的可解释性与任务进展对齐度。
原文摘要 · Abstract (English)
Pretrained-feature world models provide a useful substrate for robot imagination, but visual or latent prediction alone does not determine whether an imagined future satisfies task-relevant predicates. Long-horizon manipulation requires progress signals that are relational, predicate-level, and physically grounded: whether an object has moved, whether a drawer or contact state has changed, whether a placement predicate is satisfied, and whether a candidate future is reliable enough for execution. We introduce \textbf{EV-WM}, a predicate-grounded verification framework for world-model planning. EV-WM rolls out candidate futures in pretrained visual-feature space, decodes them into structured event states, and scores them using task-progress, semantic-consistency, physical-feasibility, and uncertainty terms. The verifier guides sampling-based planning, gates candidate actions, and, in the contact-sensitive LIBERO wine-rack setting, selects among PPO-generated proposals. Across navigation, deformable-object, wall-constrained, and language-described manipulation studies, EV-WM shows that predicate-grounded verification can make feature-space world-model planning more interpretable and better aligned with task progress.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。