让机器人通过物理感知的时间流学习执行历史,提升长程操作成功率。
TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation

- 用物理对齐的时序流监督学习紧凑的历史表征
- 在LIBERO长任务上达96.6%成功率,多阶段操作优势明显
- 无需部署时计算几何或运动估计,适合实时机器人控制
视觉-语言-动作(VLA)模型利用预训练的视觉-语言表征进行机器人控制,但简单添加历史帧难以可靠捕捉近期物理变化。这在多阶段操作中尤为严重,因视觉相似状态可能需不同动作取决于先前执行。为此,我们提出TemporalFlow-VLA,通过物理对齐的时序监督学习紧凑的执行历史。利用记录的机器人状态、机器人几何与标定相机,构建仅用于训练的机器人-表面时序流作为监督目标,并对两个执行对齐的时序查询进行监督,为动作专家提供结构化历史信息。几何监督路径不参与部署。TemporalFlow-VLA在LIBERO上平均成功率达97.63% ± 0.26%,其中LIBERO Long任务为96.60% ± 0.87%;在12个RoboTwin任务中,清洁/随机设置下分别达到85.5%和84.2%的成功率。其优势在长程、多阶段操作中最为显著。可控历史干预显示,动作预测依赖于历史内容与时间顺序。异步特征缓存下,时序条件保持单帧级服务器采样延迟,无额外历史编码开销。总体而言,TemporalFlow-VLA提供了无需部署时显式运动估计或几何处理的紧凑、物理对齐的有序执行历史接口。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models leverage pretrained vision-language representations for robot control, yet simply adding historical frames does not reliably capture recent physical change. This is especially problematic in multi-stage manipulation, where visually similar states may require different actions depending on prior execution. To address this challenge, we present TemporalFlow-VLA, which learns compact execution history through physically grounded temporal supervision. Using recorded robot states, robot geometry, and calibrated cameras, we construct robot-surface temporal flow as a training-only target and supervise two execution-aligned temporal queries that provide structured history to the action expert. The geometric supervision path is not evaluated at deployment. TemporalFlow-VLA achieves 97.63 +/- 0.26% average success on LIBERO, including 96.60 +/- 0.87% on LIBERO Long, and 85.5%/84.2% Clean/Randomized success across 12 RoboTwin tasks. It shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation. Controlled history interventions show that action prediction depends on both historical content and temporal order. With asynchronous feature caching, temporal conditioning maintains single-frame-level server-side sampling latency without additional historical-encoding overhead. Overall, TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。