arXiv:2603.07691cs.ROcs.CV2026-03中稿 · ICRA被引 1

让机器人从人类示范中学习物体操作姿势与接触点的联合预测。

RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation

  • 以操作指令为条件,联合预测接触区域和对应姿势。
  • 在真实机器人上实现跨任务、跨类别的强泛化能力。
  • 适合需要精准操作的机器人抓取与交互场景。

理解空间可操作性——即物体交互的接触区域及其对应的接触姿势——对机器人有效操作物体至关重要。现有方法多仅关注接触区域定位,将姿势估计独立处理,易因区域与姿势不一致导致任务失败。本文提出RoboPCA,一种以姿势为中心的可操作性预测框架,能根据任务指令联合预测合适的接触区域与姿势。为支持规模化数据采集,我们设计了Human2Afford数据整理流程,通过人类示范自动恢复场景级3D信息并推断姿势中心的可操作性标注。该流程利用场景深度与物体掩码提供3D上下文与定位信息,通过跟踪接触区域内的物体点并分析手物交互模式,建立从3D手部网格到机器人末端执行器朝向的映射。RoboPCA融合几何-外观线索(通过RGB-D编码器)并引入掩码增强特征以突出任务相关物体区域,集成于基于扩散模型的框架中,在图像数据集、仿真环境及真实机器人上均优于基线方法,展现出强大的跨任务与跨类别泛化能力。

原文摘要 · Abstract (English)

Understanding spatial affordances -- comprising the contact regions of object interaction and the corresponding contact poses -- is essential for robots to effectively manipulate objects and accomplish diverse tasks. However, existing spatial affordance prediction methods mainly focus on locating the contact regions while delegating the pose to independent pose estimation approaches, which can lead to task failures due to inconsistencies between predicted contact regions and candidate poses. In this work, we propose RoboPCA, a pose-centered affordance prediction framework that jointly predicts task-appropriate contact regions and poses conditioned on instructions. To enable scalable data collection for pose-centered affordance learning, we devise Human2Afford, a data curation pipeline that automatically recovers scene-level 3D information and infers pose-centered affordance annotations from human demonstrations. With Human2Afford, scene depth and the interaction object's mask are extracted to provide 3D context and object localization, while pose-centered affordance annotations are obtained by tracking object points within the contact region and analyzing hand-object interaction patterns to establish a mapping from the 3D hand mesh to the robot end-effector orientation. By integrating geometry-appearance cues through an RGB-D encoder and incorporating mask-enhanced features to emphasize task-relevant object regions into the diffusion-based framework, RoboPCA outperforms baseline methods on image datasets, simulation, and real robots, and exhibits strong generalization across tasks and categories.

机器人操作可操作性预测姿态估计人类示范

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。