提出可保持空间变换一致性的机器人操作模型,提升复杂场景泛化能力。
EquAct: An SE(3)-Equivariant Multi-Task Transformer for Open-Loop Robotic Manipulation
- 采用球面傅里叶特征的点云U-net实现三维空间等变推理
- 在18个仿真任务上达到当前最优性能,物理实验也表现优异
- 适合需要强空间泛化的多任务机器人控制场景
Transformer架构能通过联合处理自然语言指令和3D观测,从示范中学习语言驱动的多任务3D开环操作策略。然而,尽管机器人策略和语言指令均蕴含丰富的三维几何结构,标准Transformer缺乏几何一致性保障,导致场景发生SE(3)变换时行为不可预测。本文将SE(3)等变性作为策略与语言共有的关键结构特性,提出新型SE(3)-等变多任务Transformer EquAct。EquAct在理论上保证SE(3)等变性,包含两个核心组件:(1) 基于点云的高效SE(3)-等变U-net,使用球面傅里叶特征进行策略推理;(2) 用于语言条件化的SE(3)-不变特征线性调制(iFiLM)层。为评估其空间泛化能力,我们在18个RLBench仿真任务上测试了SE(3)和SE(2)场景扰动,并在4个真实物理任务上验证。EquAct在所有仿真与物理任务上均达到领先水平。
原文摘要 · Abstract (English)
Transformer architectures can effectively learn language-conditioned, multi-task 3D open-loop manipulation policies from demonstrations by jointly processing natural language instructions and 3D observations. However, although both the robot policy and language instructions inherently encode rich 3D geometric structures, standard transformers lack built-in guarantees of geometric consistency, often resulting in unpredictable behavior under SE(3) transformations of the scene. In this paper, we leverage SE(3) equivariance as a key structural property shared by both policy and language, and propose EquAct-a novel SE(3)-equivariant multi-task transformer. EquAct is theoretically guaranteed to be SE(3) equivariant and consists of two key components: (1) an efficient SE(3)-equivariant point cloud-based U-net with spherical Fourier features for policy reasoning, and (2) SE(3)-invariant Feature-wise Linear Modulation (iFiLM) layers for language conditioning. To evaluate its spatial generalization ability, we benchmark EquAct on 18 RLBench simulation tasks with both SE(3) and SE(2) scene perturbations, and on 4 physical tasks. EquAct performs state-of-the-art across these simulation and physical tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。