arXiv:2409.14891cs.ROcs.CV2024-09被引 15

让机器人主动移动摄像头看清楚物体,再决定怎么抓。

Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation

  • 相机和夹爪分步决策,动态调整视角与抓取姿态
  • 在RLBench的8个任务中,成功率显著高于基线方法
  • 适合摄像头视角受限的真实机械臂操作场景

在真实场景中,许多机器人操作任务因遮挡和视野受限而受阻,被动观测模型依赖固定或腕装摄像头难以应对。本文研究视觉受限下的机器人操作问题,提出一种任务驱动的异步主动视觉-动作模型。该模型串联相机的下一步最佳视角(NBV)策略与夹爪的下一步最佳姿态(NBP)策略,在少样本强化学习框架下进行传感器-运动协调训练。该方法使智能体能根据任务目标主动调整第三人称摄像头视角,进而推断合适的操作动作。我们在RLBench的8个视角受限任务上进行了训练与评估,结果表明本模型始终优于基线算法,展现出在处理视觉约束下的操作任务中的有效性。

原文摘要 · Abstract (English)

In real-world scenarios, many robotic manipulation tasks are hindered by occlusions and limited fields of view, posing significant challenges for passive observation-based models that rely on fixed or wrist-mounted cameras. In this paper, we investigate the problem of robotic manipulation under limited visual observation and propose a task-driven asynchronous active vision-action model.Our model serially connects a camera Next-Best-View (NBV) policy with a gripper Next-Best Pose (NBP) policy, and trains them in a sensor-motor coordination framework using few-shot reinforcement learning. This approach allows the agent to adjust a third-person camera to actively observe the environment based on the task goal, and subsequently infer the appropriate manipulation actions.We trained and evaluated our model on 8 viewpoint-constrained tasks in RLBench. The results demonstrate that our model consistently outperforms baseline algorithms, showcasing its effectiveness in handling visual constraints in manipulation tasks.

机器人操作主动视觉强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。