arXiv:2503.02405cs.RO2025-03

用3D场景信息比纯视觉图像能显著提升机械臂抓取性能

A comparison of visual representations for real-world reinforcement learning in the context of vacuum gripping

  • 对比了视觉与3D空间编码器在强化学习中的表现
  • 3D输入使抓取成功率显著高于纯视觉方案
  • 适合关注真实机器人控制与感知融合的研究者

在真实世界物体操作中,需依赖传感器反馈构建响应式决策策略。本研究探究不同编码器在强化学习框架中解析机器人臂周围局部环境空间信息的能力。基于SERL框架,我们构建了一个样本高效且稳定的训练基础,同时保持训练时间短。实验在真空夹爪的盒体抓取任务上进行,结果表明:包含空间信息的3D输入显著优于纯视觉输入。代码与评估视频见https://github.com/nisutte/voxel-serl。

原文摘要 · Abstract (English)

When manipulating objects in the real world, we need reactive feedback policies that take into account sensor information to inform decisions. This study aims to determine how different encoders can be used in a reinforcement learning (RL) framework to interpret the spatial environment in the local surroundings of a robot arm. Our investigation focuses on comparing real-world vision with 3D scene inputs, exploring new architectures in the process. We built on the SERL framework, providing us with a sample efficient and stable RL foundation we could build upon, while keeping training times minimal. The results of this study indicate that spatial information helps to significantly outperform the visual counterpart, tested on a box picking task with a vacuum gripper. The code and videos of the evaluations are available at https://github.com/nisutte/voxel-serl.

强化学习3D感知机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。