用单目相机像素变化直接感知物体运动,实现敏捷抓取的仿真到现实迁移。
Pixel2Catch: Multi-Agent Sim-to-Real Transfer for Agile Manipulation with a Single RGB Camera
- 基于单帧图像的像素级变化识别物体运动,避免3D位置估计
- 设计异构多智能体框架,机械臂与多指手协同训练并成功迁移到真实世界
- 适合高自由度机器人敏捷操作任务,尤其关注视觉驱动的实时控制
要接住一个被抛出的物体,机器人必须能够及时感知物体的运动并生成控制动作。本文提出一种新方法,不显式估计物体的3D位置,而是通过单个RGB图像提取的像素级视觉信息来识别物体运动。这些视觉线索捕捉了物体位置和尺度的变化,使策略能够推理其运动状态。此外,为在由带多指手的机械臂构成的高自由度系统中实现稳定学习,我们设计了一个异构多智能体强化学习框架,将机械臂和手视为具有不同角色的独立智能体。每个智能体使用特定的角色观测和奖励进行协作训练,所学策略成功从仿真环境迁移至真实世界。
原文摘要 · Abstract (English)
To catch a thrown object, a robot must be able to perceive the object's motion and generate control actions in a timely manner. Rather than explicitly estimating the object's 3D position, this work focuses on a novel approach that recognizes object motion using pixel-level visual information extracted from a single RGB image. Such visual cues capture changes in the object's position and scale, allowing the policy to reason about the object's motion. Furthermore, to achieve stable learning in a high-DoF system composed of a robot arm equipped with a multi-fingered hand, we design a heterogeneous multi-agent reinforcement learning framework that defines the arm and hand as independent agents with distinct roles. Each agent is trained cooperatively using role-specific observations and rewards, and the learned policies are successfully transferred from simulation to the real world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。