通过解耦观察与动作的因果关系,提升机器人模仿学习的泛化能力。
Improving Generalization Ability of Robotic Imitation Learning by Resolving Causal Confusion in Observations
- 基于干预机制显式学习观测与专家动作的因果关系。
- 在ALOHA双臂仿真中显著改善复杂模仿学习的泛化性能。
- 方法可无缝嵌入现有架构,适合复杂机械臂任务部署。
近期模仿学习在机器人操作领域取得显著进展,但现有技术在面对轻微环境变化时仍存在泛化能力差的问题。本文旨在增强复杂模仿学习算法在训练与部署环境不一致情况下的泛化能力。为避免无关观测带来的混淆,提出显式学习观测成分与专家动作之间的因果关系,借鉴[6]的框架,通过干预模仿学习策略来学习因果结构函数。我们从理论上阐明,在机器人操作的复杂模仿学习过程中,无需像[6]那样对图像输入特征进行解耦,该要求并非必要。因此,提出一种简单且可嵌入当前主流架构(如Action Chunking Transformer [31])的因果结构学习框架。在Mujoco中的ALOHA [31] 双臂机器人仿真环境中验证,该方法能有效缓解现有复杂模仿学习算法的泛化问题。
原文摘要 · Abstract (English)
Recent developments in imitation learning have considerably advanced robotic manipulation. However, current techniques in imitation learning can suffer from poor generalization, limiting performance even under relatively minor domain shifts. In this work, we aim to enhance the generalization capabilities of complex imitation learning algorithms to handle unpredictable changes from the training environments to deployment environments. To avoid confusion caused by observations that are not relevant to the target task, we propose to explicitly learn the causal relationship between observation components and expert actions, employing a framework similar to [6], where a causal structural function is learned by intervention on the imitation learning policy. Disentangling the feature representation from image input as in [6] is hard to satisfy in complex imitation learning process in robotic manipulation, we theoretically clarify that this requirement is not necessary in causal relationship learning. Therefore, we propose a simple causal structure learning framework that can be easily embedded in recent imitation learning architectures, such as the Action Chunking Transformer [31]. We demonstrate our approach using a simulation of the ALOHA [31] bimanual robot arms in Mujoco, and show that the method can considerably mitigate the generalization problem of existing complex imitation learning algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。