系统研究机械臂动作空间设计,发现合理设计能显著提升学习效果
Demystifying Action Space Design for Robotic Manipulation Policies
- 从时间与空间维度分析动作空间设计,量化其对学习的影响
- 预测动作增量(delta)比绝对值表示平均提升性能,任务空间更利于泛化
- 双臂机器人实测1.3万次,结果可指导实际机械臂策略设计
动作空间的设计在基于模仿的机器人抓取策略学习中起关键作用,从根本上影响策略学习的优化路径。尽管近期研究集中于扩大训练数据和模型容量,动作空间的选择仍依赖经验或历史设计,导致对机器人策略设计的理解模糊。为此,我们开展大规模系统性实验,验证动作空间对机器人策略学习具有显著且复杂的影响。我们从时间和空间两个维度拆解动作设计空间,系统分析其对策略可学习性和控制稳定性的影响。基于在双臂机器人上进行的13,000+次真实世界试运行,以及对4个场景下500多个训练模型的评估,我们比较了绝对值与增量表示、关节空间与任务空间参数化的优劣。大规模结果表明:合理设计为预测增量动作可一致提升性能;关节空间更利于控制稳定,任务空间则更利于泛化。
原文摘要 · Abstract (English)
The specification of the action space plays a pivotal role in imitation-based robotic manipulation policy learning, fundamentally shaping the optimization landscape of policy learning. While recent advances have focused heavily on scaling training data and model capacity, the choice of action space remains guided by ad-hoc heuristics or legacy designs, leading to an ambiguous understanding of robotic policy design philosophies. To address this ambiguity, we conducted a large-scale and systematic empirical study, confirming that the action space does have significant and complex impacts on robotic policy learning. We dissect the action design space along temporal and spatial axes, facilitating a structured analysis of how these choices govern both policy learnability and control stability. Based on 13,000+ real-world rollouts on a bimanual robot and evaluation on 500+ trained models over four scenarios, we examine the trade-offs between absolute vs. delta representations, and joint-space vs. task-space parameterizations. Our large-scale results suggest that properly designing the policy to predict delta actions consistently improves performance, while joint-space and task-space representations offer complementary strengths, favoring control stability and generalization, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。