arXiv:2605.21133cs.RO2026-05被引 1

用双模块框架让机器人在复杂环境里自主完成全身操作任务。

Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum

论文配图:Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum
图 1 · 摘自论文原文
  • 分两步:先主动感知空间,再生成通用动作
  • 无需真实机器人数据就能跨任务泛化,实测表现强
  • 适合做具身智能、机器人操作系统的研发者参考

本文研究复杂3D环境中人形机器人全身操作任务。相比桌面场景,该任务面临两大挑战:1)复杂环境中空间关系多样,理解困难;2)动作生成难以泛化,因真实机器人数据少且成本高。为此,我们提出一种可泛化的拟人化运动-操作框架,利用多智能体大模型的空间感知与动作生成能力。框架包含两个组件:主动空间脑(Active Spatial Brain)负责主动感知环境并规划任务与子任务分解;通用动作小脑(Generalizable Action Cerebellum)根据前一模块决策生成可执行机器人动作,无需任务特定真实数据。为评估框架,我们设计了一组空间操作任务,从空间理解能力和真实机器人性能两方面进行测试。结果表明,该框架在多种任务与环境中均表现出色。

原文摘要 · Abstract (English)

In this paper, we explore spatial-aware humanoid whole-body manipulation task. Compared with tabletop settings, this task poses two key challenges: 1) Spatial understanding is challenging in complex 3D environments with diverse spatial relations. 2) Action generation is difficult to generalize, as limited and costly real-robot data restricts data-driven models generalization. To address these challenges, we propose a generalizable humanoid loco-manipulation framework that leverages the spatial perception and action generation capabilities of multi-agent large models. Specifically, our framework includes two components: Active Spatial Brain for active spatial perception and decision-making, and Generalizable Action Cerebellum for executable robot action generation. The first component actively perceives the spatial scene and makes decisions on task planning and subtask decomposition. The second component generate executable robot actions based on the decisions made by the first module without needs of task-specific real robot data. To benchmark our framework, we design a set of spatial manipulation tasks from two perspectives: evaluating spatial perception and understanding, and assessing real-robot task performance. The results demonstrate strong performance on both aspects across diverse tasks and environments.

人形机器人空间感知动作生成泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。