arXiv:2602.02741cs.RO2026-02被引 2

从一次人类操作中学习物体关节模型,无需先验知识。

PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations

  • 仅用一次人类示范的点云序列,端到端估计关节参数与操作顺序。
  • 在多种物体上提升27%以上的关节轴和状态估计准确率。
  • 适用于真实场景,能恢复交互中才显露的被遮挡关节。

关节建模使机器人能够学习可动物体的关节参数,从而有效执行操作任务,并可用于下游技能学习或规划。现有方法通常依赖于对物体的先验知识,如关节数量或类型;部分方法无法恢复仅在交互中显现的被遮挡关节;另一些则需要每件物体的大量多视角图像,这在真实环境中不切实际。此外,以往工作忽略了操作顺序的重要性,而这对多自由度物体(如洗碗机)至关重要。本文提出PokeNet,一种端到端框架,仅需一次人类示范即可从未知物体的点云序列中估计关节模型,推断操作顺序并追踪关节状态。PokeNet在多种物体上表现优于现有最优方法,平均提升关节轴与状态估计准确率超过27%,并在仿真与真实环境均验证了其有效性。

原文摘要 · Abstract (English)

Articulation modeling enables robots to learn joint parameters of articulated objects for effective manipulation which can then be used downstream for skill learning or planning. Existing approaches often rely on prior knowledge about the objects, such as the number or type of joints. Some of these approaches also fail to recover occluded joints that are only revealed during interaction. Others require large numbers of multi-view images for every object, which is impractical in real-world settings. Furthermore, prior works neglect the order of manipulations, which is essential for many multi-DoF objects where one joint must be operated before another, such as a dishwasher. We introduce PokeNet, an end-to-end framework that estimates articulation models from a single human demonstration without prior object knowledge. Given a sequence of point cloud observations of a human manipulating an unknown object, PokeNet predicts joint parameters, infers manipulation order, and tracks joint states over time. PokeNet outperforms existing state-of-the-art methods, improving joint axis and state estimation accuracy by an average of over 27% across diverse objects, including novel and unseen categories. We demonstrate these gains in both simulation and real-world environments.

关节建模人演示机器人点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。