通过人工标注关键点提升机器人策略的泛化能力
P3-PO: Prescriptive Point Priors for Visuo-Spatial Generalization of Robot Policies
- 用人工标注的语义点构建环境状态表示,替代直接视觉输入
- 在真实任务中相比旧方法提升43%成功率,新物体场景提升80%
- 适合需要强泛化能力的机器人操作研究者使用
开发能在多变环境和不同物体下稳健运行的通用机器人策略仍是机器人学习中的核心挑战。尽管已有大量工作聚焦于收集大规模机器人数据集并设计相应策略架构,但直接从视觉输入学习常导致策略脆弱,难以超越训练数据分布。本文提出一种名为P3-PO(Prescriptive Point Priors for Policies)的新框架,利用计算机视觉与机器人学习的最新进展,构建独特的环境状态表示,以提升机器人操作策略的分布外泛化能力。该表示通过两步获得:首先由人类标注单个演示帧上的若干语义有意义的点;随后利用现成的视觉模型将这些点传播至整个数据集。生成的点作为输入用于最先进的策略架构进行学习。在四个真实世界任务上的实验表明,在与训练环境相同的设置下,相比先前方法整体提升了43%的性能。此外,针对新物体实例的任务中提升达58%,在更杂乱环境中提升达80%。机器人表现视频可在 point-priors.github.io 查看。
原文摘要 · Abstract (English)
Developing generalizable robot policies that can robustly handle varied environmental conditions and object instances remains a fundamental challenge in robot learning. While considerable efforts have focused on collecting large robot datasets and developing policy architectures to learn from such data, naively learning from visual inputs often results in brittle policies that fail to transfer beyond the training data. This work presents Prescriptive Point Priors for Policies or P3-PO, a novel framework that constructs a unique state representation of the environment leveraging recent advances in computer vision and robot learning to achieve improved out-of-distribution generalization for robot manipulation. This representation is obtained through two steps. First, a human annotator prescribes a set of semantically meaningful points on a single demonstration frame. These points are then propagated through the dataset using off-the-shelf vision models. The derived points serve as an input to state-of-the-art policy architectures for policy learning. Our experiments across four real-world tasks demonstrate an overall 43% absolute improvement over prior methods when evaluated in identical settings as training. Further, P3-PO exhibits 58% and 80% gains across tasks for new object instances and more cluttered environments respectively. Videos illustrating the robot's performance are best viewed at point-priors.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。