用图像平面轨迹学习机器人抓取,提升泛化能力与鲁棒性
Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation

- 将3D操作转为2D图像空间轨迹预测,降低学习难度
- 通过三角测量无损恢复末端执行器位置,支持多视角融合
- 利用等变增强提升对相机变化的适应性,适合真实场景部署
将操纵动作表示为相机平面上的2D轨迹,为学习复杂3D操纵策略提供紧凑且可解释的基础。然而,这也带来了轨迹出框和精度受限的问题。我们提出Pix2Act,一种模仿学习方法,通过在每个相机平面生成连续的图像空间关键点轨迹,并利用三角测量无损恢复末端执行器姿态,将高维3D控制问题重构为更简单、更易学习的2D预测任务。关键在于,它使观测与动作处于同一坐标空间,支持等变变换——同时旋转各相机图像及其对应的图像空间动作。我们分析了该增强的对称性,并设计了能融合多视角且尊重各自旋转的网络架构。结果,Pix2Act隐式扩大了数据分布的支持范围,学习到跨变换的不变动作结构,显著提升泛化能力和整体性能。在多种模拟与真实世界操纵任务中,Pix2Act优于当前最优基线,且对相机扰动保持鲁棒。
原文摘要 · Abstract (English)
Representing manipulation actions as 2D trajectories in the camera plane provides a compact and interpretable basis for learning complex 3D manipulation policies. However, it also creates challenges from out-of-frame trajectories and limited precision. We propose Pix2Act, an imitation learning method that addresses these challenges by generating continuous image-space keypoint trajectories in each camera plane and losslessly recovering end-effector poses via triangulation. This reformulates high-dimensional 3D control as a simpler, more learnable 2D prediction problem. Crucially, it aligns observations and actions in the same coordinate space, enabling equivariant transformations to jointly rotate individual camera images together with their image-space actions. We analyze the symmetry properties of this augmentation and design a network architecture that can fuse multiple camera views while respecting their per-view rotations. As a result, Pix2Act implicitly enlarges the support of the data distribution and learns invariant action structures across transformations, yielding improved generalization and overall performance. Across diverse simulated and real-world manipulation tasks, Pix2Act outperforms state-of-the-art baselines and remains robust under camera perturbations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。