提出可泛化的双手非抓握操作原语,无需强化学习即可实现跨机器人迁移。
BiNoMaP: Learning Category-Level Bimanual Non-Prehensile Manipulation Primitives
- 从第一视角视频中提取双手动作轨迹,经几何感知优化生成可执行操作原语。
- 通过物体尺寸等几何属性参数化,实现对未见物体的类别级泛化。
- 支持不同机械结构的双臂机器人直接复用,无需重新设计技能结构。
非抓握操作(如推、戳、撬动、缠绕)因接触频繁且解析困难而研究不足。本文提出双臂通用配置下的双手非抓握操作原语(BiNoMaP),突破单臂或依赖环境支撑的局限。采用三阶段无强化学习框架:首先从第一视角视频中提取双手动作轨迹;其次引入几何感知后优化算法,消除感知噪声与形态差异,生成符合预设运动模式的可执行原语;最后通过物体尺寸等几何属性参数化,实现对未见物体的类别级泛化。重要的是,该方法支持跨机器人平台迁移,同一原语可在两种不同运动学构型的真实双臂机器人上部署,无需重构技能结构。大量真实机器人实验在多样物体与空间配置下验证了方法的有效性、高效性及强泛化能力。
原文摘要 · Abstract (English)
Non-prehensile manipulation, encompassing ungraspable actions such as pushing, poking, pivoting, and wrapping, remains underexplored due to its contact-rich and analytically intractable nature. We revisit this problem from two perspectives. First, instead of relying on single-arm setups or favorable environmental supports (e.g., walls or edges), we advocate a generalizable dual-arm configuration and establish a suite of Bimanual Non-prehensile Manipulation Primitives (BiNoMaP). Second, departing from prevailing RL-based approaches, we propose a three-stage, RL-free framework for learning structured non-prehensile skills. We begin by extracting bimanual hand motion trajectories from egocentric video demonstrations. Since these coarse trajectories suffer from perceptual noise and morphological discrepancies, we introduce a geometry-aware post-optimization algorithm to refine them into executable manipulation primitives consistent with predefined motion patterns. To enable category-level generalization, the learned primitives are further parameterized by object-relevant geometric attributes, primarily size, allowing adaptation to unseen instances with significant shape variations. Importantly, BiNoMaP supports cross-embodiment transfer: the same primitives can be deployed on two real-world dual-arm platforms with distinct kinematic configurations, without redesigning skill structures. Extensive real-robot experiments across diverse objects and spatial configurations demonstrate the effectiveness, efficiency, and strong generalization capability of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。