arXiv:2602.12978cs.ROcs.AI2026-02中稿 · Robotics: Science …被引 17

让视觉语言动作模型在分段执行时更平滑,减少卡顿和误切换。

Learning Native Continuation for Action Chunking Flow Policies

  • 训练时用混合动作与噪声初始化去噪过程,模拟部分动作信息。
  • 通过重塑流形动态,使推理时每步引导保持一致性,提升平滑性。
  • 支持不同推理延迟,适合真实场景中的机器人操控任务。

动作分段使视觉语言动作(VLA)模型能实时运行,但直接分段执行常在分段边界产生不连续。实时分段(RTC)虽缓解此问题,但作为外部机制,导致多模态误切换且轨迹不够内在平滑。本文提出Legato,一种面向动作分段流模型策略的训练时连续性学习方法。Legato从已知动作与噪声的调度混合分布初始化去噪过程,使模型暴露于部分动作信息;同时重塑学习到的流动力学,确保在每步引导下训练与推理的一致性。此外,训练中引入随机调度条件,以支持不同推理延迟并实现可控平滑度。实证表明,Legato生成更平滑的轨迹,减少误切换,降低犹豫时间,缩短任务完成时长。大量真实世界实验显示,其在五个操作任务上持续优于RTC,轨迹平滑度与任务完成时间分别提升约10%。

原文摘要 · Abstract (English)

Action chunking enables Vision Language Action (VLA) models to run in real time, but naive chunked execution often exhibits discontinuities at chunk boundaries. Real-Time Chunking (RTC) alleviates this issue but is external to the policy, leading to spurious multimodal switching and trajectories that are not intrinsically smooth. We propose Legato, a training-time continuation method for action-chunked flow-based VLA policies. Specifically, Legato initializes denoising from a schedule-shaped mixture of known actions and noise, exposing the model to partial action information. Moreover, Legato reshapes the learned flow dynamics to ensure that the denoising process remains consistent between training and inference under per-step guidance. Legato further uses randomized schedule condition during training to support varying inference delays and achieve controllable smoothness. Empirically, Legato produces smoother trajectories and reduces spurious multimodal switching during execution, leading to less hesitation and shorter task completion time. Extensive real-world experiments show that Legato consistently outperforms RTC across five manipulation tasks, achieving approximately 10% improvements in both trajectory smoothness and task completion time.

动作分段流模型机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。