arXiv:2411.16959cs.ROcs.AI2024-11ICRA被引 18

通过因果不变性与几何对称性增强机器人模仿学习的数据效率

RoCoDA: Counterfactual Data Augmentation for Data-Efficient Robot Learning from Demonstrations

  • 融合因果不变性与SE(3)对称性生成合成示范数据
  • 在5个任务中提升泛化能力与样本效率,减少对真实数据依赖
  • 适合研究机器人模仿学习、数据高效强化学习的学者

机器人模仿学习面临环境复杂性和数据采集成本高的挑战。本文提出RoCoDA,将不变性、等变性和因果性统一于一个框架中,用于增强模仿学习的数据增强。通过修改环境状态中与任务无关的部分而不改变策略输出,实现因果不变性;同时利用SE(3)等变性对物体位姿施加刚体变换,并调整对应动作生成合成示范。在五个机器人操作任务上进行广泛实验,结果表明,相比现有最优数据增强方法,RoCoDA显著提升了策略性能、泛化能力和样本效率。策略能稳健泛化至未见过的物体姿态、纹理及干扰物。此外,观察到重抓取等涌现行为,表明使用RoCoDA训练的策略对任务动态有更深层理解。通过结合不变性、等变性与因果性,RoCoDA为模仿学习中的数据增强提供了原则性方法,弥合了几何对称性与因果推理之间的差距。

原文摘要 · Abstract (English)

Imitation learning in robotics faces significant challenges in generalization due to the complexity of robotic environments and the high cost of data collection. We introduce RoCoDA, a novel method that unifies the concepts of invariance, equivariance, and causality within a single framework to enhance data augmentation for imitation learning. RoCoDA leverages causal invariance by modifying task-irrelevant subsets of the environment state without affecting the policy's output. Simultaneously, we exploit SE(3) equivariance by applying rigid body transformations to object poses and adjusting corresponding actions to generate synthetic demonstrations. We validate RoCoDA through extensive experiments on five robotic manipulation tasks, demonstrating improvements in policy performance, generalization, and sample efficiency compared to state-of-the-art data augmentation methods. Our policies exhibit robust generalization to unseen object poses, textures, and the presence of distractors. Furthermore, we observe emergent behavior such as re-grasping, indicating policies trained with RoCoDA possess a deeper understanding of task dynamics. By leveraging invariance, equivariance, and causality, RoCoDA provides a principled approach to data augmentation in imitation learning, bridging the gap between geometric symmetries and causal reasoning. Project Page: https://rocoda.github.io

机器人学习数据增强模仿学习因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。