arXiv:2606.08288cs.RO2026-06被引 1

让机器人模型记住运动轨迹而非零散帧,提升长程操作的稳定性和流畅性。

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

论文配图:MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model
图 1 · 摘自论文原文
  • 用连续轨迹场编码历史运动,替代独立帧的拼接
  • 在模拟和真实机器人上实现更平滑、更直接的长程操作
  • 适合需要精准运动控制的具身智能任务

视觉-语言-动作(VLA)模型越来越依赖历史、深度或4D特征来解决长程操作中的歧义。然而,更多时空证据并不一定更好:若注入的证据缺乏运动一致性,可能引入几何漂移、时间线索断裂和动作生成不稳定。这引发一个核心问题:VLA应记住过去帧,还是记住连接它们的运动?我们提出MotionVLA,一种运动历史接口,将短时仅历史视频窗口转换为紧凑、连续的轨迹场标记。不同于将历史视为独立提取的稀疏帧,MotionVLA将近期观测表示为物理一致的运动证据。当前视觉标记通过查询该历史来获取任务相关的运动信息,并在轨迹引导监督下重新整合进VLA流。在多个仿真基准和初步真实机器人滚动测试中,MotionVLA提升了长程操作性能,同时生成更平滑、更直接的执行动作。结果表明,有效的VLA记忆不仅在于提供更多的4D上下文,更在于暴露可用于控制的运动一致证据。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiotemporal evidence is not necessarily better: when the injected evidence is not motion-consistent, it can introduce geometric drift, fragmented temporal cues, and unstable action generation. This raises a simple question: should a VLA remember past frames, or remember the motion that connects them? We introduce MotionVLA, a motion-history interface that converts a short past-only video window into compact, time-continuous trajectory-field tokens. Instead of treating history as a sparse set of ndependently lifted frames, MotionVLA represents recent observations as physically coherent motion evidence. Current visual tokens query this history to retrieve task-relevant motion information, which is then recoupled into the VLA stream under trajectory-grounded supervision. Experiments across simulation benchmarks and preliminary real-robot rollouts show that MotionVLA improves long-horizon manipulation while producing smoother and more direct executions. These results suggest that effective VLA memory is not just about providing more 4D context, but about exposing motion-consistent evidence that is usable for control.

视觉-语言-动作机器人控制运动建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。