arXiv:2606.31493cs.RO2026-06被引 1

用统一时间流建模过去现在未来,提升机器人抓取的泛化能力

ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning

论文配图:ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning
图 1 · 摘自论文原文
  • 基于稀疏3D关键点构建时空统一表征,融合物体与机械臂的动态信息
  • 在14个仿真和5个真实任务中表现优于主流扩散策略模型
  • 特别适合长序列、非马尔可夫的复杂操作场景,如连续抓取

视觉信号在策略学习中至关重要,能帮助模型捕捉物体运动与交互动态。如同人类通过过往经验与预期结果推理行为,有效的策略应整合历史交互与未来预测。然而,现有视觉动作策略通常孤立建模历史或未来动态,缺乏统一的时间表征。本文提出ChronoFlow,一种通过物体与夹爪的稀疏3D关键点捕捉过去、当前与未来交互动态的时序统一表征。基于此,我们设计ChronoFlow-Policy,一种基于扩散模型的视觉动作策略,通过联合训练目标同时学习ChronoFlow与动作序列。在14个仿真任务与5个真实世界操作任务上的实验表明,ChronoFlow-Policy持续优于强基准扩散策略,并在长时程与非马尔可夫操作场景中显著提升鲁棒性。

原文摘要 · Abstract (English)

Visual signals play a crucial role in policy learning by enabling models to capture object motion and interaction dynamics. Just as humans reason about actions using both past experience and anticipated outcomes, effective policies should integrate past interactions with future predictions. However, existing visuomotor policies typically model either historical context or future dynamics in isolation, lacking a unified temporal representation of interaction dynamics. In this work, we introduce ChronoFlow, a temporally unified representation that captures past, current, and future interaction dynamics through sparse 3D keypoints of both objects and the gripper. Based on this representation, we propose ChronoFlow-Policy, a diffusion-based visuomotor policy that jointly learns ChronoFlow and action sequences through a co-training objective. Experiments on 14 simulated tasks and 5 real-world manipulation tasks demonstrate that ChronoFlow-Policy consistently outperforms strong diffusion-policy baselines and improves robustness in long-horizon and non-Markovian manipulation scenarios. Our project page is available at https://the-kamisato-sii.github.io/ChronoFlow-Policy-project-page/.

视觉动作策略扩散模型时序建模机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。