arXiv:2606.21148cs.RO2026-06

让机器人在任意朝向下精准抓握咖啡杯把手,无需重新训练。

Pose-Agnostic Robotic Functional Grasping via Observation-Action Canonicalization

论文配图:Pose-Agnostic Robotic Functional Grasping via Observation-Action Canonicalization
图 1 · 摘自论文原文
  • 将视觉与动作统一到物体中心坐标系,实现姿态无关的抓取策略。
  • 仿真中对正反放置的杯子成功率超93%,真实机器人零样本迁移达80%。
  • 适合需要高泛化能力的工业抓取场景,尤其对细长把手物体有效。

功能性机器人抓取需在多样物体几何形状与姿态下保持任务特异性接触精度。本文以杯柄抓取为例,研究细长把手、实例差异及正/倒置摆放带来的感知与控制敏感性。传统开环抓取检测方法易受细柄结构估计误差影响;学习型视觉-动作策略需隐式处理外观与动作方向耦合变化,限制泛化能力。为此提出AnyMug框架,一种基于观察-动作规范化(observation-action canonicalization)的闭环强化学习方法,完全在仿真中训练单一策略,并实现零样本部署于真实机器人。该方法将深度观测与末端执行器动作统一映射至共享物体重心坐标系,使策略始终看到一致的杯子视角并输出规范动作方向,从而跨姿态复用同一抓取行为。引入手柄感知奖励,促进精确接近、夹爪对齐与对指定位;结合姿态课程与领域随机化,提升训练稳定性和仿真到现实的迁移效果。在仿真中,AnyMug对未见过的正/倒置杯子成功率均超过93%;零样本迁移到真实Franka Panda机械臂,在5个未见实物杯子上成功率达80%。

原文摘要 · Abstract (English)

Functional robotic grasping requires a policy that generalizes across diverse object geometries and poses while maintaining task-specific contact precision. We study this challenge through mug-handle grasping, where thin handles, instance variation, and upright or inverted placements make both perception and control sensitive to object configuration. Grasp pose detection methods operate open-loop and are sensitive to estimation errors on thin handle structures. Learned visuomotor policies must implicitly learn to handle the coupled variation in visual appearance and action direction induced by different object placements, limiting generalization. We propose AnyMug, a canonicalized visuomotor reinforcement learning framework for functional grasping that trains a single closed-loop policy entirely in simulation and deploys it zero-shot on a real robot. AnyMug introduces observation-action canonicalization, which transforms both the depth observation and the predicted end-effector action into a shared object-centric frame. The policy therefore sees a consistent mug-centered view and emits actions in a canonical direction regardless of mug placement, allowing the same grasping behavior to be reused across configurations. A handle-aware reward further encourages precise approach, gripper alignment, and opposing-finger placement, while a pose curriculum and domain randomization improve training stability and sim-to-real transfer. In simulation, AnyMug achieves over 93% success rate on both unseen upright and inverted mugs and transfers zero-shot to a real Franka Panda, reaching 80% success rate on 5 held-out physical mugs across both pose categories.

机器人抓取强化学习姿态不变仿真迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。