通过视觉逆强化学习让机器人模仿人类操作,更自然地完成抓取和倒液任务。
Visual IRL for Human-Like Robotic Manipulation
- 用RGB-D关键点直接做状态特征,通过逆强化学习学人类动作奖励函数。
- 新神经符号模型将人手动作映射到机械臂,保持自然运动轨迹。
- 适合希望快速部署人机协作的制造业场景,提升机器人自然度。
我们提出一种新型方法,使协作机器人(cobots)能通过观察人类动作学习并以类人方式执行操纵任务。该方法属于“从观察中学习”(LfO)范式,相比从零编程,可更快集成到工业场景。我们引入视觉逆强化学习(Visual IRL),直接使用观察视频中每帧的RGB-D关键点作为状态特征,输入逆强化学习(IRL)模型,学习出将关键点映射为奖励值的奖励函数。该奖励函数通过一种新型神经符号动力学模型,从人类动作映射至机器人手臂,实现末端执行器位置相似的同时最小化关节调整,旨在保留人类动作的自然动态特性。与以往仅关注末端位置的方法不同,本方法将人体多个关节角度映射到对应机器人关节,并利用逆运动学模型最小调整关节角,以精确实现末端定位。我们在两个真实操纵任务上评估该方法:第一个是农产品处理任务,涉及捡起、检查并按是否受损分类放置洋葱;第二个是液体倾倒任务,包括拿起瓶子、倒入指定容器并丢弃空瓶。结果表明,该方法显著提升了机器人操作的类人程度,增强了制造场景中人机协作的兼容性。
原文摘要 · Abstract (English)
We present a novel method for collaborative robots (cobots) to learn manipulation tasks and perform them in a human-like manner. Our method falls under the learn-from-observation (LfO) paradigm, where robots learn to perform tasks by observing human actions, which facilitates quicker integration into industrial settings compared to programming from scratch. We introduce Visual IRL that uses the RGB-D keypoints in each frame of the observed human task performance directly as state features, which are input to inverse reinforcement learning (IRL). The inversely learned reward function, which maps keypoints to reward values, is transferred from the human to the cobot using a novel neuro-symbolic dynamics model, which maps human kinematics to the cobot arm. This model allows similar end-effector positioning while minimizing joint adjustments, aiming to preserve the natural dynamics of human motion in robotic manipulation. In contrast with previous techniques that focus on end-effector placement only, our method maps multiple joint angles of the human arm to the corresponding cobot joints. Moreover, it uses an inverse kinematics model to then minimally adjust the joint angles, for accurate end-effector positioning. We evaluate the performance of this approach on two different realistic manipulation tasks. The first task is produce processing, which involves picking, inspecting, and placing onions based on whether they are blemished. The second task is liquid pouring, where the robot picks up bottles, pours the contents into designated containers, and disposes of the empty bottles. Our results demonstrate advances in human-like robotic manipulation, leading to more human-robot compatibility in manufacturing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。