arXiv:2606.02486cs.RO2026-06被引 2

让视觉语言动作模型提前预测动态物体位置,提升抓取成功率

Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation

论文配图:Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation
图 1 · 摘自论文原文
  • 用运动感知的隐空间世界模型预测未来视觉特征
  • 动态场景下成功率从58%提升至97%,物理机器人测试中多数任务全胜
  • 无需训练主模型,仅加490万参数即可增强现有系统

视觉-语言-动作(VLA)模型在静态操作中表现良好,但在物体运动时失效。它们将当前观测映射为动作,并假设观察与执行间场景静止,导致物体移动过快时延迟超出可抓取时间。本文提出AHEAD(前瞻视界外推自适应动力学),一种增强冻结VLA模型的预测-执行框架。一个小规模世界模型基于操作视频训练,预测VLA特征空间中的未来图像块,条件依赖于光流计算的每块速度与加速度。语言与运动显著性掩码聚焦于任务相关区域,模型沿自适应时间窗滚动预测,直至不确定性超过阈值。冻结的动作解码器接收预测的未来特征而非当前特征。AHEAD为70亿参数的OpenVLA增加490万参数,在20个动态仿真场景中成功率达79%~97%,最强基线仅为31%~58%。在真实UFactory xArm 7机器人上,对传送带、滚球任务成功率分别为29/30至30/30,击打球任务23/30,投掷物捕捉任务19/30,所有基线均为0/30。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models generalize across static manipulation but fail when objects move during task execution. They map the current observation to an action and assume the scene is stationary between observation and execution, so at any non-trivial object speed the resulting latency exceeds the time available to grasp. We close this gap with AHEAD (Anticipatory Horizon Extrapolation with Adaptive Dynamics), a predict-then-act wrapper that augments a frozen VLA with a motion-aware latent world model. A small world model trained on manipulation video forecasts future patch tokens in the VLA's feature space, conditioned on per-token velocity and acceleration from optical flow. A language-and-motion saliency mask concentrates prediction on task-relevant patches, and the model rolls forward for an adaptive horizon, halting when prediction uncertainty crosses a threshold. The frozen action decoder then receives the predicted future tokens in place of the current ones. AHEAD adds 4.9M parameters to a frozen 7B OpenVLA and reaches 79 to 97% success across 20 dynamic simulation scenarios where the strongest baseline reaches 31 to 58%. On a physical UFactory xArm 7, AHEAD succeeds on 29/30 to 30/30 on three conveyor and rolling-ball tasks, 23/30 on paddle interception, and 19/30 on projectile catching where every baseline scores 0/30.

VLA模型动态抓取预测控制机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。