发现机器人长序列开环执行主要为应对非马尔可夫示范,长上下文策略更优。
Revisiting Open-Loop Execution in Robotics: Toward Reactive, Higher-Performing Policies

- 通过长上下文感知,让策略更好模仿非马尔可夫式示范动作
- 在6个任务中验证:专家非马尔可夫性比累积误差影响更大
- 长上下文下闭环策略反应最快,性能超越开环执行
动作分块——即预测一串动作并开环执行前缀——已成为机器人操作模仿学习的重要技术。但长期开环执行会降低系统反应能力,限制纠错能力。现有研究认为其优势源于缓解误差累积、吸收推理延迟或平滑动作,但缺乏严格验证。本文表明,长开环执行主要帮助短上下文策略模仿‘非马尔可夫性’的专家示范。在四个仿真和两个真实世界任务中,我们发现专家非马尔可夫性显著影响任务成功率与开环时长的关系。进一步分析误差累积的影响,结果显示尽管其存在,但远不如非马尔可夫性重要。最后,当策略拥有足够长上下文时,开环不再有益,最实时的闭环策略表现最佳。因此,长上下文、高反应性策略应成为更合理且高性能的新范式。
原文摘要 · Abstract (English)
Action chunking --- the practice of predicting a sequence of actions and executing a prefix open-loop --- has emerged as a key enabler of recent progress in imitation learning for robotic manipulation. However, executing long open-loop prefixes reduces reactivity, limiting policies' ability to correct for errors. Further, the mechanisms underlying these performance benefits remain poorly understood: prior works cite mitigating compounding errors, absorbing inference latency, or smoothing motions, but provide limited controlled evidence or guidance for preserving reactivity. In this work, we argue that long open-loop execution primarily helps short-context policies imitate "non-Markovian demonstrations". Across four simulation and two real-world tasks, we show that expert non-Markovianity strongly shapes the relationship between task success and open-loop execution horizon. Further, we investigate the impact of compounding errors --- the prevailing explanation for long open-loop execution in prior work --- and find that while they matter, expert non-Markovianity has a much stronger impact in our experimental setting. Finally, we show that when policies are provided with a sufficiently long context, open-loop execution is no longer beneficial and the most reactive, closed-loop policies perform best. While imitation learning has seen great success using long open-loop execution, our findings motivate long-context, reactive policies as a more principled and performant paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。