通过轨迹引导与世界模型校准,提升机器人移动操作的精准度和成功率。
DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation

- 用统一模型同时预测末端执行器轨迹和全身动作,显式引导基臂协同
- 在真实任务中平均成功率提升至90.0%,接触密集任务提升最显著
- 适合需要高精度移动操作的工业或服务机器人场景
移动操作需在不断变化的视角和接触条件下协调底盘与机械臂运动,其动作空间远大于固定基座操作。现有视觉-语言-动作(VLA)策略存在两方面局限:(i) 直接将观测映射为全身动作块,在未显式规划任务空间路径的情况下搜索大动作空间,导致基臂协同预测不精确;(ii) 预测动作开环执行,无法验证动作是否实现预期运动,导致控制误差与未建模接触累积,造成计划与实际运动偏差。本文提出 DreamTrajectory,一种轨迹引导的语言条件移动操作框架,针对每项局限引入一个组件。针对(i),DreamTrajectory 在单一动作专家中联合预测意图级末端执行器轨迹与全身动作块,使轨迹显式指导基臂动作生成而非隐含。针对(ii),引入轻量级轨迹世界模型,预测候选动作块所引发的实际轨迹,并通过测试时搜索-预测-评分流程选择与计划轨迹对齐最优的动作。在 MS-HAB 上,轨迹引导使平均成功率从 32.3% 提升至 47.5%,测试时优化进一步达 54.8%,尤其在接触丰富的刚性物体任务中增益最大。在三个真实移动操作任务中,对应平均成功率为 63.3%、81.7% 和 90.0%。
原文摘要 · Abstract (English)
Mobile manipulation requires a robot to coordinate base and arm motion under continuously changing viewpoints and contact conditions, within an action space far larger than that of fixed-base manipulation. Existing Vision-Language-Action (VLA) policies are limited in two respects. (i)They map observations directly to whole-body action chunks, searching this large action space without an explicit task-space motion plan, which makes coordinated base--arm prediction imprecise. (ii)They execute the predicted chunk open-loop, without checking whether the actions can realize the motion the policy intended, so control errors and unmodeled contacts accumulate into a gap between planned and realized motion. We present DreamTrajectory, a trajectory-guided framework for language-conditioned mobile manipulation that introduces one component for each limitation. Addressing(i), DreamTrajectory jointly predicts an intention-level end-effector trajectory and a whole-body action chunk in a single action expert, so that the trajectory explicitly guides base--arm action generation instead of remaining implicit. Addressing(ii), a lightweight trajectory world model predicts the trajectory that a candidate action chunk would induce, and a test-time search--predict--score procedure selects the candidate best aligned with the planned trajectory. On MS-HAB, trajectory guidance raises average success from 32.3% to 47.5% and test-time refinement further to 54.8%, with the largest gains on contact-rich articulated-object tasks. On three real-world mobile manipulation tasks, the corresponding average success rates are 63.3%, 81.7%, and 90.0%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。