用隐变量反馈控制器建模技能习得,提升机器人学习的鲁棒性。
A Probabilistic Model for Skill Acquisition with Switching Latent Feedback Controllers
- 将混合密度网络重解释为基于隐状态的反馈控制器库。
- 在人类示范下训练,任务成功率显著提升且抗观测噪声能力增强。
- 适合需要高鲁棒性技能学习的机器人部署场景。
操作任务通常由多个子任务构成,每个子任务代表一种特定技能。掌握这些技能对机器人实现自主性、效率、适应性和环境协作至关重要。从示范中学习可使机器人快速获取新技能,而无需从零开始,示范通常按顺序组合技能以完成任务。行为克隆方法常依赖混合密度网络输出头来预测机器人动作。本文首次将混合密度网络重新解释为基于隐状态的反馈控制器(或技能)库,这一见解源于发现单层线性网络在功能上等同于经典反馈控制器,其权重对应控制器增益。基于此,我们构建了一个概率图模型,将技能习得过程描述为隐空间中的分段,每个技能策略在该隐空间中表现为一个反馈控制律。该方法不仅显著提升了任务成功率,还增强了对观测噪声的鲁棒性。物理机器人实验进一步表明,这种诱导出的鲁棒性有效提升了模型在真实机器人的部署表现。
原文摘要 · Abstract (English)
Manipulation tasks often consist of subtasks, each representing a distinct skill. Mastering these skills is essential for robots, as it enhances their autonomy, efficiency, adaptability, and ability to work in their environment. Learning from demonstrations allows robots to rapidly acquire new skills without starting from scratch, with demonstrations typically sequencing skills to achieve tasks. Behaviour cloning approaches to learning from demonstration commonly rely on mixture density network output heads to predict robot actions. In this work, we first reinterpret the mixture density network as a library of feedback controllers (or skills) conditioned on latent states. This arises from the observation that a one-layer linear network is functionally equivalent to a classical feedback controller, with network weights corresponding to controller gains. We use this insight to derive a probabilistic graphical model that combines these elements, describing the skill acquisition process as segmentation in a latent space, where each skill policy functions as a feedback control law in this latent space. Our approach significantly improves not only task success rate, but also robustness to observation noise when trained with human demonstrations. Our physical robot experiments further show that the induced robustness improves model deployment on robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。