让机器人通过智能探索快速识别物体属性,提升复杂任务成功率。
Poke and Strike: Learning Task-Informed Exploration Policies
- 用强化学习训练探索策略,奖励来自任务对属性误差的敏感性。
- 90%任务成功率,平均探索时间低于1.2秒,远超基线方法。
- 适合需要自主感知与调整的动态机器人任务,如物理操作。
在许多动态机器人任务中,例如将冰球击入可达工作空间外的球门,机器人必须先识别物体的关键物理属性才能成功执行任务,否则无法自我修复或重试而需人工干预。为此,我们提出一种基于强化学习的任务引导探索方法,利用特权任务策略对属性估计误差的敏感性自动生成奖励信号来训练探索策略。同时引入基于不确定性的机制,判断何时从探索转入任务执行,确保属性估计精度的同时最小化探索时间。该方法在击打任务上实现90%的成功率,平均探索时间低于1.2秒,显著优于基线方法(最高仅40%成功率),且无需测试时重复仿真查询与重训练。此外,我们验证了任务引导奖励能捕捉不同任务中物理属性的相对重要性,涵盖击打任务和经典CartPole示例。最后,我们在KUKA iiwa机械臂的物理平台上验证了该方法可有效识别物体属性并自适应调整任务执行。
原文摘要 · Abstract (English)
In many dynamic robotic tasks, such as striking pucks into a goal outside the reachable workspace, the robot must first identify the relevant physical properties of the object for successful task execution, as it is unable to recover from failure or retry without human intervention. To address this challenge, we propose a task-informed exploration approach, based on reinforcement learning, that trains an exploration policy using rewards automatically generated from the sensitivity of a privileged task policy to errors in estimated properties. We also introduce an uncertainty-based mechanism to determine when to transition from exploration to task execution, ensuring sufficient property estimation accuracy with minimal exploration time. Our method achieves a 90% success rate on the striking task with an average exploration time under 1.2 seconds, significantly outperforming baselines that achieve at most 40% success or require inefficient querying and retraining in a simulator at test time. Additionally, we demonstrate that our task-informed rewards capture the relative importance of physical properties in both the striking task and the classical CartPole example. Finally, we validate our approach by demonstrating its ability to identify object properties and adjust task execution in a physical setup using the KUKA iiwa robot arm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。