用几何先验与脉冲网络提升人形机器人任务泛化与数据效率。
Generalizable Geometric Prior and Recurrent Spiking Feature Learning for Humanoid Robot Manipulation
- 结合2D几何先验增强视觉语言模型的3D场景理解能力。
- 在多个真实机器人平台实现比现有方法更优的泛化性能。
- 适合研究机器人感知-决策一体化与高效学习的学者。
人形机器人操作是执行多样化人类级任务的关键研究方向,涉及高层语义推理与底层动作生成。然而,精确场景理解及从人类示范中实现样本高效学习仍是重大挑战,严重制约现有框架的适用性与泛化能力。本文提出一种新型递归几何先验多模态策略(RGMP-S),支持高层技能推理与数据高效的动作合成。为将高层推理锚定在物理现实,我们引入轻量级2D几何归纳偏置,使视觉语言模型实现精准3D场景理解。具体地,构建长时程几何先验技能选择器,有效对齐语义指令与空间约束,在未见环境中实现鲁棒泛化。针对动作生成的数据效率问题,提出递归自适应脉冲网络,通过递归脉冲参数化机器人-物体交互,保持时空一致性,充分提取长时程动态特征,并缓解稀疏示范下的过拟合问题。在Maniskill仿真基准及三个异构真实机器人系统(定制人形机器人、桌面机械臂、商用机器人平台)上开展大量实验,结果表明该方法显著优于现有最先进基线,验证了各模块在多种泛化场景中的有效性。为促进可复现性,源代码与视频演示已公开于https://github.com/xtli12/RGMP-S.git。
原文摘要 · Abstract (English)
Humanoid robot manipulation is a crucial research area for executing diverse human-level tasks, involving high-level semantic reasoning and low-level action generation. However, precise scene understanding and sample-efficient learning from human demonstrations remain critical challenges, severely hindering the applicability and generalizability of existing frameworks. This paper presents a novel RGMP-S, Recurrent Geometric-prior Multimodal Policy with Spiking features, facilitating both high-level skill reasoning and data-efficient motion synthesis. To ground high-level reasoning in physical reality, we leverage lightweight 2D geometric inductive biases to enable precise 3D scene understanding within the vision-language model. Specifically, we construct a Long-horizon Geometric Prior Skill Selector that effectively aligns the semantic instructions with spatial constraints, ultimately achieving robust generalization in unseen environments. For the data efficiency issue in robotic action generation, we introduce a Recursive Adaptive Spiking Network. We parameterize robot-object interactions via recursive spiking for spatiotemporal consistency, fully distilling long-horizon dynamic features while mitigating the overfitting issue in sparse demonstration scenarios. Extensive experiments are conducted across the Maniskill simulation benchmark and three heterogeneous real-world robotic systems, encompassing a custom-developed humanoid, a desktop manipulator, and a commercial robotic platform. Empirical results substantiate the superiority of our method over state-of-the-art baselines and validate the efficacy of the proposed modules in diverse generalization scenarios. To facilitate reproducibility, the source code and video demonstrations are publicly available at https://github.com/xtli12/RGMP-S.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。