arXiv:2607.18840cs.RO2026-07被引 2

让机器人更懂任务进展,能根据指令自主规划并精细执行。

WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory

论文配图:WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory
图 1 · 摘自论文原文
  • 用记忆增强的模型同时存短期视觉和长期事件历史,提升任务理解。
  • 在500万段动作数据上预训练,实现从指令到动作的精准对齐。
  • 支持文本、图像或视频提示控制,适合长程复杂任务场景。

世界动作模型(WAMs)通过联合建模视觉状态变化与机器人动作,为机械臂操作提供了新范式。然而现有模型受限于有限的时间上下文、粗粒度的整段语言监督以及主要依赖文本的条件输入,难以追踪任务进展、实现细粒度的语言-视频-动作对齐,并限制了视觉上下文推理与跨平台迁移。本文提出WorldScape Policy 2.0,一种可调控的世界动作模型,其融合因果短时视觉记忆与增强推理的长短期记忆机制。短时记忆以DiT预填充形式保留近期观测,维持局部交互动态;长时记忆将历史视觉语言模型输出组织为全局历史、局部活跃与事件边界三类表征,实现进度感知的检索。检索到的历史信息增强感知并生成自回归规划令牌,隐含子目标条件以支持自主规划;语义强制进一步将事件级指令语义注入该潜在规划路径。为实现细粒度多模态可控性,我们构建了ManipEvent-5M数据集,包含近500万段带对齐动作轨迹、整段任务指令、片段级子任务描述、目标图像和视频演示的事件片段。这些设计统一了从高层指令自主规划与从细粒度文本、目标图像或视频上下文可控执行的接口。仿真与真实平台实验均表明,该模型在长周期自主规划、细粒度指令遵循及上下文适应方面表现优异。

原文摘要 · Abstract (English)

World Action Models (WAMs) offer a promising paradigm for robotic manipulation by jointly modeling visual state transitions and robot actions. However, existing WAMs are constrained by limited temporal context, coarse episode-level language supervision, and predominantly text-only conditioning, which hinder task-progress tracking and fine-grained language-video-action grounding while limiting visual-context reasoning and cross-embodiment transfer. In this paper, we introduce WorldScape Policy 2.0, a controllable WAM with reasoning-augmented long short-term memory. Its causal short-term visual memory supplies recent observations as DiT prefill to preserve local interaction dynamics, while its long short-term event memory organizes historical VLM outputs into global-history, local-active, and event-boundary representations for progress-aware retrieval. The retrieved history augments perception and autoregressively generated planning tokens, yielding an implicit subgoal condition for autonomous planning; semantic forcing further transfers event-level instruction semantics into this latent planning pathway. To establish fine-grained multimodal controllability, we construct ManipEvent-5M, an event-grounded embodied pretraining dataset containing nearly 5 million event segments with aligned action trajectories, episode-level task instructions, segment-level subtask captions, goal images, and video demonstrations. These designs provide a unified interface for autonomous planning from high-level instructions and controllable execution from fine-grained text, goal-image, or video-context prompts. Experiments in both simulation and real-world platforms demonstrate superior capabilities in long-horizon autonomous planning, fine-grained instruction following and in-context adaptation.

机器人动作建模多模态控制长程规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。