控制器增益影响机器人学习效果,应按学习方法选而非按任务需求定。
Tune to Learn: How Controller Gains Shape Robot Policy Learning
- 根据学习范式调整位置控制器增益,而非仅依据任务刚度需求。
- 行为克隆在柔顺且过阻尼增益下表现最好,强化学习则对增益不敏感。
- 仿真到现实迁移在刚性或过阻尼增益下受损,需避免此类设置。
位置控制器已成为执行学习到的抓取策略的主要接口。然而一个关键的设计决策仍缺乏研究:如何为策略学习选择控制器增益?传统观点是根据期望的任务柔顺性或刚度来设定增益。但在状态条件策略与控制器结合时,有效刚度来自学习反应与控制动态的相互作用,而非仅由增益决定。本文主张增益选择应以可学习性为导向:即不同增益设置对所用学习算法的适应程度。我们系统研究了位置控制器增益对现代机器人学习流水线三个核心组件的影响:行为克隆、从零开始的强化学习以及仿真到现实的迁移。通过多个任务和机器人本体的广泛实验发现:(1) 行为克隆在柔顺且过阻尼增益配置下获益最大;(2) 强化学习在合理超参数调优下可在所有增益配置中成功;(3) 仿真到现实迁移在刚性和过阻尼增益下受到损害。这些结果表明,最优增益选择不应基于目标任务行为,而应取决于所采用的学习范式。
原文摘要 · Abstract (English)
Position controllers have become the dominant interface for executing learned manipulation policies. Yet a critical design decision remains understudied: how should we choose controller gains for policy learning? The conventional wisdom is to select gains based on desired task compliance or stiffness. However, this logic breaks down when controllers are paired with state-conditioned policies: effective stiffness emerges from the interplay between learned reactions and control dynamics, not from gains alone. We argue that gain selection should instead be guided by learnability: how amenable different gain settings are to the learning algorithm in use. In this work, we systematically investigate how position controller gains affect three core components of modern robot learning pipelines: behavior cloning, reinforcement learning from scratch, and sim-to-real transfer. Through extensive experiments across multiple tasks and robot embodiments, we find that: (1) behavior cloning benefits from compliant and overdamped gain regimes, (2) reinforcement learning can succeed across all gain regimes given compatible hyperparameter tuning, and (3) sim-to-real transfer is harmed by stiff and overdamped gain regimes. These findings reveal that optimal gain selection depends not on the desired task behavior, but on the learning paradigm employed. Project website: https://younghyopark.me/tune-to-learn
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。