arXiv:2608.26821cs.RO2026-08

让机器人通过物理感知的时间流学习执行历史,提升长程操作成功率。

TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation

论文配图:TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation
图 1 · 摘自论文原文
  • 用物理对齐的时序流监督学习紧凑的历史表征
  • 在LIBERO长任务上达96.6%成功率,多阶段操作优势明显
  • 无需部署时计算几何或运动估计,适合实时机器人控制

视觉-语言-动作(VLA)模型利用预训练的视觉-语言表征进行机器人控制,但简单添加历史帧难以可靠捕捉近期物理变化。这在多阶段操作中尤为严重,因视觉相似状态可能需不同动作取决于先前执行。为此,我们提出TemporalFlow-VLA,通过物理对齐的时序监督学习紧凑的执行历史。利用记录的机器人状态、机器人几何与标定相机,构建仅用于训练的机器人-表面时序流作为监督目标,并对两个执行对齐的时序查询进行监督,为动作专家提供结构化历史信息。几何监督路径不参与部署。TemporalFlow-VLA在LIBERO上平均成功率达97.63% ± 0.26%,其中LIBERO Long任务为96.60% ± 0.87%;在12个RoboTwin任务中,清洁/随机设置下分别达到85.5%和84.2%的成功率。其优势在长程、多阶段操作中最为显著。可控历史干预显示,动作预测依赖于历史内容与时间顺序。异步特征缓存下,时序条件保持单帧级服务器采样延迟,无额外历史编码开销。总体而言,TemporalFlow-VLA提供了无需部署时显式运动估计或几何处理的紧凑、物理对齐的有序执行历史接口。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models leverage pretrained vision-language representations for robot control, yet simply adding historical frames does not reliably capture recent physical change. This is especially problematic in multi-stage manipulation, where visually similar states may require different actions depending on prior execution. To address this challenge, we present TemporalFlow-VLA, which learns compact execution history through physically grounded temporal supervision. Using recorded robot states, robot geometry, and calibrated cameras, we construct robot-surface temporal flow as a training-only target and supervise two execution-aligned temporal queries that provide structured history to the action expert. The geometric supervision path is not evaluated at deployment. TemporalFlow-VLA achieves 97.63 +/- 0.26% average success on LIBERO, including 96.60 +/- 0.87% on LIBERO Long, and 85.5%/84.2% Clean/Randomized success across 12 RoboTwin tasks. It shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation. Controlled history interventions show that action prediction depends on both historical content and temporal order. With asynchronous feature caching, temporal conditioning maintains single-frame-level server-side sampling latency without additional historical-encoding overhead. Overall, TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment.

机器人操作时序建模视觉-语言-动作长程任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。