arXiv:2502.01773cs.ROcs.CV2025-02被引 3

利用对称性提升机器人抓取任务的样本效率

Coarse-to-Fine 3D Keyframe Transporter

  • 基于跨帧特征相关性设计可迁移的抓取策略
  • 在仿真任务上平均性能优于基线10%以上,实物实验提升55%
  • 适合需要高效泛化能力的机器人操作研究者

近期关键帧模仿学习(Keyframe IL)进展使基于学习的智能体能够解决多种操作任务。然而,多数方法忽略了问题中的丰富对称性,导致样本效率低下。本文识别并利用了关键帧动作方案中的双等变对称性,设计出能泛化至工作空间和抓取物体变换的策略。主要贡献有二:第一,分析了关键帧动作方案的双等变性质,提出一种源自Transporter Networks的键帧传输器,通过抓取物与场景特征间的交叉相关性评估动作;第二,提出一种计算高效的粗到精SE(3)动作评估方案,用于推理平移与旋转的耦合动作。该方法在广泛仿真任务中平均性能优于强基线超过10%,在4项物理实验中平均提升55%。

原文摘要 · Abstract (English)

Recent advances in Keyframe Imitation Learning (IL) have enabled learning-based agents to solve a diverse range of manipulation tasks. However, most approaches ignore the rich symmetries in the problem setting and, as a consequence, are sample-inefficient. This work identifies and utilizes the bi-equivariant symmetry within Keyframe IL to design a policy that generalizes to transformations of both the workspace and the objects grasped by the gripper. We make two main contributions: First, we analyze the bi-equivariance properties of the keyframe action scheme and propose a Keyframe Transporter derived from the Transporter Networks, which evaluates actions using cross-correlation between the features of the grasped object and the features of the scene. Second, we propose a computationally efficient coarse-to-fine SE(3) action evaluation scheme for reasoning the intertwined translation and rotation action. The resulting method outperforms strong Keyframe IL baselines by an average of >10% on a wide range of simulation tasks, and by an average of 55% in 4 physical experiments.

模仿学习机器人操作对称性3D抓取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。