让机器人提前看懂物体运动,提升复杂操作的鲁棒性
OFlow: Injecting Object-Aware Temporal Flow Matching for Robust Robotic Manipulation

- 用统一语义空间同时预测未来和识别物体
- 在多个基准上成功率提升,尤其在分布外场景表现更好
- 适合需要精准动作控制的机器人任务
可靠机器人操作不仅需要预测场景随时间的变化,还需在复杂场景中识别任务相关物体。现有视觉语言模型通常仅处理当前帧,未来预测与物体感知常在独立隐空间中学习。本文提出OFlow(将物体感知时间流匹配注入视觉语言模型),通过共享语义隐空间统一时间前瞻与物体感知。该方法使用时间流匹配预测未来隐状态,将其分解为强调物理相关线索、过滤无关变化的物体感知表示,并基于此生成连续动作。将OFlow集成至视觉语言模型流程后,显著提升在分布外情况下的控制可靠性。在LIBERO、LIBERO-Plus、MetaWorld及SimplerEnv等基准与真实任务上的实验表明,物体感知前瞻性可持续增强鲁棒性与成功率。
原文摘要 · Abstract (English)
Robust robotic manipulation requires not only predicting how the scene evolves over time, but also recognizing task-relevant objects in complex scenes. However, existing VLA models face two limitations. They typically act only on the current frame, while future prediction and object-aware reasoning are often learned in separate latent spaces. We propose OFlow (injecting Object-Aware Temporal Flow Matching into VLAs), a framework that addresses both limitations by unifying temporal foresight and object-aware reasoning in a shared semantic latent space. Our method forecasts future latents with temporal flow matching, factorizes them into object-aware representations that emphasize physically relevant cues while filtering task-irrelevant variation, and conditions continuous action generation on these predictions. By integrating OFlow into VLA pipelines, our method enables more reliable control under distribution shifts. Extensive experiments across LIBERO, LIBERO-Plus, MetaWorld, and SimplerEnv benchmarks and real-world tasks demonstrate that object-aware foresight consistently enhances robustness and success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。