arXiv:2606.05773cs.RO2026-06被引 1

让机器人在想象中闭环试错,提升策略评估准确性

PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation

论文配图:PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation
图 1 · 摘自论文原文
  • 构建分段世界模型,支持视觉-语言-动作策略的闭环推理
  • 真实任务中评估误差从63.2%降至12.0%,接近真实表现
  • 可学习成功与失败轨迹,提升模拟执行的真实性

视觉-语言-动作(VLA)策略在真实机器人任务中以闭环方式运行:机器人观察环境,执行一个动作片段,并根据前一动作产生的观测结果决定下一步。然而,现有世界模型多限于对预采集动作轨迹进行开环预测,无法支持闭环评估,即每一步动作需基于前序动作生成的观测进行条件判断。为此,我们提出PiL-World,一种专为策略闭环评估设计的分段式世界模型。给定当前观测和由VLA策略生成的动作轨迹,PiL-World生成与策略回放一致的多视角未来观测,匹配策略所需的图像输入。通过交替进行VLA推理与世界模型预测,实现无需每步真实执行的闭环评估。为提升回放保真度,PiL-World结合动作驱动的视觉控制(来自俯视视角机器人运动)与编码任务上下文的潜在历史,联合预测互补的多视角观测。除成功示范外,还从失败轨迹中学习,使模拟回放更贴近真实策略执行分布。我们在三个真实双臂操作任务上评估了PiL-World,结果显示其生成的想象回放高度一致于真实机器人执行。更重要的是,相较于基线,其在闭环世界模型评估中估计的VLA成功率与真实世界回放测量值之间的误差从63.2%降至12.0%。

原文摘要 · Abstract (English)

Vision-language-action (VLA) policies operate in a closed loop in real-world robot tasks: a robot observes the scene, executes an action chunk, and conditions its next decision on the resulting observation. However, most existing world models for robot action evaluation are limited to open-loop prediction along pre-collected action trajectories. This prevents them from supporting closed-loop VLA evaluation, where each action chunk must be conditioned on the observation generated by the previous execution. To address this gap, we propose PiL-World, a chunk-wise world model designed for policy-in-the-loop VLA evaluation. Given the current observation and the action trajectory rolled out by a VLA policy, PiL-World generates multi-view future observations that are consistent with the VLA rollout and match the image inputs required by the policy. By alternating between VLA inference and world-model prediction, PiL-World enables closed-loop evaluation without real robot execution at every step. To improve rollout fidelity, PiL-World conditions video generation on action-derived visual control from head-view robot motion and latent histories that encode task execution context, while jointly predicting complementary multi-view observations. Beyond successful teleoperated demonstrations, it also learns from failed execution trajectories, helping the imagined rollouts better match the distribution of real policy executions. We evaluate PiL-World on three real dual-arm manipulation tasks. PiL-World generates imagined rollouts that are highly consistent with real robot executions. More importantly, compared with the baseline, it reduces the error between VLA success rates measured in real-world rollouts and those estimated through closed-loop world-model evaluation from 63.2% to 12.0%.

世界模型闭环评估机器人决策模拟推演

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。