用几何先验提升人形机器人操作的泛化能力与数据效率
RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation
- 结合几何先验与视觉语言模型,自动生成适应新场景的操作序列
- 在未见过场景中实现87%任务成功率,训练数据效率提升5倍
- 适合关注机器人泛化能力与少样本学习的研究者
人形机器人具备执行多样化人类级技能的潜力。然而,当前研究主要依赖数据驱动方法,需大量训练数据才能实现鲁棒的多模态决策与可泛化的视觉运动控制。这些方法忽视了未知场景中的几何推理,且对机器人-目标关系建模效率低下,造成训练资源浪费。为此,我们提出端到端的递归几何先验多模态策略(RGMP),融合几何-语义技能推理与数据高效的视觉运动控制。感知方面,提出几何先验技能选择器,将几何归纳偏置注入视觉语言模型,仅需少量空间常识调优即可生成适应新场景的技能序列。为实现数据高效运动生成,引入自适应递归高斯网络,将机器人-物体交互参数化为紧凑的高斯过程层级,递归编码多尺度空间关系,在稀疏示范下仍能生成灵巧动作。在人形机器人和桌面双臂机器人上评估,RGMP在泛化测试中达到87%任务成功率,数据效率比最先进模型高出5倍,展现出由几何-语义推理与递归高斯适配带来的卓越跨域泛化能力。
原文摘要 · Abstract (English)
Humanoid robots exhibit significant potential in executing diverse human-level skills. However, current research predominantly relies on data-driven approaches that necessitate extensive training datasets to achieve robust multimodal decision-making capabilities and generalizable visuomotor control. These methods raise concerns due to the neglect of geometric reasoning in unseen scenarios and the inefficient modeling of robot-target relationships within the training data, resulting in significant waste of training resources. To address these limitations, we present the Recurrent Geometric-prior Multimodal Policy (RGMP), an end-to-end framework that unifies geometric-semantic skill reasoning with data-efficient visuomotor control. For perception capabilities, we propose the Geometric-prior Skill Selector, which infuses geometric inductive biases into a vision language model, producing adaptive skill sequences for unseen scenes with minimal spatial common sense tuning. To achieve data-efficient robotic motion synthesis, we introduce the Adaptive Recursive Gaussian Network, which parameterizes robot-object interactions as a compact hierarchy of Gaussian processes that recursively encode multi-scale spatial relationships, yielding dexterous, data-efficient motion synthesis even from sparse demonstrations. Evaluated on both our humanoid robot and desktop dual-arm robot, the RGMP framework achieves 87% task success in generalization tests and exhibits 5x greater data efficiency than the state-of-the-art model. This performance underscores its superior cross-domain generalization, enabled by geometric-semantic reasoning and recursive-Gaussion adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。