仅用人类演示视频训练机器人,无需实操数据。
Point Policy: Unifying Observations and Actions with Key Points for Robot Manipulation
- 用视觉模型将人手动作转为机械臂动作,关键点捕捉物体状态
- 8个真实任务中性能比之前方法提升75%,新物体上提升74%
- 适合想低成本训练通用机器人的研究者和工程师
构建能在多种环境和物体类型中操作的机器人智能体仍面临重大挑战,通常需要大量数据采集。这在机器人领域尤为受限,因每个数据点都需在真实世界中物理执行。因此,亟需替代性数据源和能从这类数据中学习的框架。本文提出 Point Policy,一种仅依赖离线人类演示视频、无需遥控操作数据的学习机器人策略方法。该方法利用先进的视觉模型和策略架构,将人类手部姿态转换为机器人姿态,并通过语义有意义的关键点捕捉物体状态,实现与形态无关的表示,从而促进有效策略学习。在8个真实任务上的实验表明,在与训练设置相同的条件下,整体性能相比先前方法提升了75%;对新物体实例的泛化能力提升74%,且对显著背景杂乱具有鲁棒性。视频演示请访问 https://point-policy.github.io/。
原文摘要 · Abstract (English)
Building robotic agents capable of operating across diverse environments and object types remains a significant challenge, often requiring extensive data collection. This is particularly restrictive in robotics, where each data point must be physically executed in the real world. Consequently, there is a critical need for alternative data sources for robotics and frameworks that enable learning from such data. In this work, we present Point Policy, a new method for learning robot policies exclusively from offline human demonstration videos and without any teleoperation data. Point Policy leverages state-of-the-art vision models and policy architectures to translate human hand poses into robot poses while capturing object states through semantically meaningful key points. This approach yields a morphology-agnostic representation that facilitates effective policy learning. Our experiments on 8 real-world tasks demonstrate an overall 75% absolute improvement over prior works when evaluated in identical settings as training. Further, Point Policy exhibits a 74% gain across tasks for novel object instances and is robust to significant background clutter. Videos of the robot are best viewed at https://point-policy.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。