arXiv:2606.07304cs.RO2026-06

让机器人通过对比动作结果来高效预测未来,提升规划精度与速度

CAPE: Contrastive Action-conditioned Parallel Encoding for Embodied Planning

论文配图:CAPE: Contrastive Action-conditioned Parallel Encoding for Embodied Planning
图 1 · 摘自论文原文
  • 用对比学习区分不同动作序列的未来影响,聚焦关键变化
  • 单次前向传播即可生成完整未来潜空间轨迹,长程预测更快
  • 在真实数据集上显著优于基线,适合需要快速决策的机器人任务

具身智能体需在执行前预测候选动作的未来后果以实现有效规划。现有视觉动态模型通过重建未来视觉状态或滚动展开密集潜在表示进行学习,导致学习能力被分散到视觉显著但与规划无关的内容上,而非驱动操作结果的动作条件变化。我们提出CAPE——一种对比式动作条件并行编码框架,通过区分不同动作序列引发的未来结果来学习视觉动态。给定初始观测和候选动作序列,CAPE可在一次前向传播中解码完整的未来潜在轨迹,并采用目标收敛对比损失训练,使相同未来结果的预测对齐,不同结果的预测分离。在真实世界数据集DROID上以及零样本迁移至RoboCasa的任务中,CAPE在未来的状态检索、离线动作匹配和闭环规划任务上均显著超越先前基线,同时在长程预测时明显降低规划推理成本。

原文摘要 · Abstract (English)

Embodied agents need to predict the future consequences of candidate actions in order to plan effectively before execution. Existing visual dynamics models learn by reconstructing future visual states or rolling out dense latent representations, which spreads learning capacity across visually salient but planning-irrelevant content rather than the action-conditioned changes that drive manipulation outcomes. We propose CAPE, a Contrastive Action-conditioned Parallel Encoding framework that learns visual dynamics by distinguishing the future outcomes induced by different action sequences. Given an initial observation and a candidate action sequence, CAPE decodes the full future latent trajectory in a single forward pass and is trained with a Goal-Convergent Contrastive Objective that aligns predictions corresponding to the same future outcome while separating those corresponding to different outcomes. On real-world DROID and zero-shot transfer to RoboCasa, CAPE substantially outperforms prior baselines on future-state retrieval, offline action matching, and closed-loop planning, while notably reducing planning-time inference cost at long prediction horizons.

具身规划视觉预测对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。