用运动信息提升机器人长序列操作的连贯性。
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
- 以运动表示替代静态图像,捕捉时间动态变化。
- 在两个基准上超越现有方法,推理延迟几乎不变。
- 适合需要长时间规划的机器人任务场景。
视觉-语言-动作(VLA)模型通过将视觉和语言线索映射为动作,推动了机器人操作的发展。然而,大多数VLA假设马尔可夫性,仅依赖当前观测,导致时间视野局限,影响长序列任务的一致性。本文提出将运动视为更紧凑、更丰富的时序上下文表征,能捕捉状态间变化并过滤静态像素噪声。基于此,我们构建了HiF-VLA(Hindsight, Insight, and Foresight for VLAs),一个以运动为中心的统一框架,支持双向时序推理:通过后见先验编码历史动态,通过前瞻推理预测未来运动,并通过后见调制联合专家融合两者,实现“边执行边思考”的长时序操作。实验表明,HiF-VLA在LIBERO-Long和CALVIN ABC-D基准上显著优于强基线,且推理延迟几乎无增加。此外,在真实世界长序列操作任务中也取得显著提升,验证了其在实际机器人场景中的广泛有效性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property, relying only on the current observation and thus suffering from temporal myopia that degrades long-horizon coherence. In this work, we view motion as a more compact and informative representation of temporal context and world dynamics, capturing inter-state changes while filtering static pixel-level noise. From this perspective, HiF-VLA equips a motion-centric world model for the VLA, enabling agents to reason about temporal dynamics for future evolution during action generation. Building on this idea, we propose HiF-VLA (Hindsight, Insight, and Foresight for VLAs), a unified framework that leverages motion for bidirectional temporal reasoning. HiF-VLA encodes past dynamics through hindsight priors, anticipates future motion via foresight reasoning, and integrates both through a hindsight-modulated joint expert to enable a ''think-while-acting'' paradigm for long-horizon manipulation. As a result, HiF-VLA surpasses strong baselines on LIBERO-Long and CALVIN ABC-D benchmarks, while incurring negligible additional inference latency. Furthermore, HiF-VLA achieves substantial improvements in real-world long-horizon manipulation tasks, demonstrating its broad effectiveness in practical robotic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。