arXiv:2608.20735cs.AIcs.RO2026-08

通过因果未来令牌蒸馏,让机械臂更准抓取传送带上的移动物体。

ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

论文配图:ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation
图 1 · 摘自论文原文
  • 从冻结的动态世界模型中蒸馏未来感知表示,保持推理时因果性
  • 真实机器人测试中抓取成功率提升12.2~22.2个百分点,快速传送带表现显著改善
  • 无需部署复杂世界模型,适合实际工业场景中的动态抓取任务

抓取运动物体需预判接触事件,但主流视觉-语言-动作(VLA)策略仅基于当前观测微调。世界动作模型(WAM)虽能学习预测动态,但部署时运行视频级教师模型或显式想象未来帧成本过高。本文提出ForeTime-VLA,一种密集型pi0.5策略,在离线阶段将当前与未来视频潜在特征压缩为64维白化目标;在线阶段,八帧历史编码器预测该目标、操作阶段及归一化过渡时间。四个未来令牌和一个阶段令牌作为VLM前缀,预测的未来状态与过渡范围则指导动作专家。训练保留原始流匹配动作目标,并加入余弦、关系几何、阶段、时间到过渡及动作等价性目标。在去重的传送带数据集上,40k步检查点在每份768个匹配窗口下测试,平均绝对误差(MAE)从0.134119降至0.130593(下降2.63%;配对自助法95%置信区间:0.82%-4.48%),L2误差降低3.02%,延迟增加2.46%-2.93%。真实机器人评估显示,静态物体抓取成功率达81.1%,慢速移动物体达58.9%,分别优于次优基线12.2和22.2个百分点。在三种传送带速度下,共完成44/90次抓取,对比pi0.5的23/90次,其中快速速度下11/30对2/30。离线姿态增益与真实机器人接触姿态失败减少的一致性支持了因果未来令牌蒸馏的有效性,实现动态抓取性能提升而不需部署世界模型教师。

原文摘要 · Abstract (English)

Manipulating moving objects requires a policy to anticipate contact events, yet vision-language-action (VLA) policies are commonly fine-tuned from the current observation alone. World action models (WAMs) learn predictive dynamics, but running a video-scale teacher or explicitly imagining future frames at deployment is costly. We introduce ForeTime-VLA, a dense pi0.5 policy that distills a future-aware, action-equivalent representation from a frozen Fast-WAM-derived teacher while remaining causal at inference. Offline, current and future video latents are compressed into a whitened 64-D target. Online, an eight-frame history encoder predicts this target together with manipulation phase and normalized time-to-transition. Four future tokens and one phase token condition the VLM prefix, while the predicted future and transition horizon condition the action expert. Training retains the original flow-matching action target and adds cosine, relational geometry, phase, time-to-transition, and action-equivalence objectives. On a deduplicated conveyor-belt dataset, we compare 40k-step checkpoints on 768 matched windows per split. Test MAE decreases from 0.134119 to 0.130593 (2.63%; paired-bootstrap 95% CI: 0.82-4.48% improvement), and test L2 decreases by 3.02%, at a 2.46-2.93% latency cost. In quantitative real-robot evaluation, ForeTime-VLA achieves 81.1% stationary and 58.9% slow-moving grasp success, exceeding the next-best reference by 12.2 and 22.2 percentage points, respectively. Across three belt speeds, it completes 44/90 grasps versus 23/90 for pi0.5, including 11/30 versus 2/30 at fast speed. The agreement between offline orientation gains and reduced real-robot contact-pose failures supports causal future-token distillation as an effective way to improve dynamic manipulation without deploying the world-model teacher.

动态抓取未来蒸馏机器人控制世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。