arXiv:2606.17833cs.RO2026-06被引 1

构建仿真基准,评估人形机器人全身动作的分层学习能力

HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning

论文配图:HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning
图 1 · 摘自论文原文
  • 分层控制:高层策略生成全身动作,底层追踪器执行
  • 7个依赖腿部协调的任务,需精确踩踏与平衡维持
  • 揭示动作表征可迁移性差,对追踪器依赖强

人形机器人有望在以人为中心环境中实现全身交互,但任务决策与全身动态执行高度耦合,导致可扩展策略学习困难。分层控制是可行方案:高层策略预测中间全身动作,底层通用运动追踪器(GMT)将其转化为稳定的人形运动。然而现有基准很少评估策略-追踪器接口本身,无法判断中间动作是否可执行、在任务分布变化下是否鲁棒,以及能否跨不同GMT后端迁移。为此,我们提出HumanoidArena,一个以仿真为基础的视角分层全身学习基准。该基准将策略学习定义为分层决策问题:高层策略基于第一人称视觉、本体感知和指令生成紧凑的全身动作,由底层GMT执行。不同于将腿部视为平面移动工具,HumanoidArena强调下肢协调在任务完成中结构必要性的交互。因此设计了7个腿部关键的物体交互(HOI)/场景交互(HSI)任务,成功需足部放置、平衡维持、姿态调整和全身重定向。为深入诊断分层系统,从扰动条件泛化和GMT条件迁移两个互补角度评估策略。实验表明,分层控制使学习策略能解决多样腿部关键交互,但性能强烈依赖追踪器,跨GMT迁移仍脆弱。这些结果使HumanoidArena成为研究可迁移中间动作表征和可扩展第一人称全身策略学习的基准。

原文摘要 · Abstract (English)

Humanoid robots promise whole-body interaction in human-centered environments, but scalable policy learning remains difficult because task-level decision-making and whole-body dynamic execution are tightly coupled. A practical solution is hierarchical control, where a high-level policy predicts intermediate whole-body actions and low-level general motion trackers (GMTs) execute them as stable humanoid motion. However, existing benchmarks rarely evaluate the policy-tracker interface itself, leaving open whether intermediate whole-body actions are executable, robust under task distribution shifts, and transferable across different GMT backends. We introduce HumanoidArena, a simulation-first benchmark for egocentric hierarchical whole-body learning. The benchmark formulates policy learning as a hierarchical decision making problem: a high-level policy converts egocentric vision, proprioception, and instructions into a compact whole-body action, which is subsequently executed by a low-level GMT. Instead of treating the legs as planar transport tools, HumanoidArena emphasizes interactions where lower-body coordination is structurally necessary in task completion. We therefore design 7 leg-critical HOI/HSI tasks in which success requires foot placement, balance maintenance, posture adjustment, and whole-body reorientation. To further diagnose the hierarchical system, we evaluate policies from two complementary perspectives: perturbation-conditioned generalization and GMT-conditioned transfer. Experiments show that hierarchical control enables learned policies to solve diverse leg-critical interactions, but performance is strongly tracker-conditioned and cross-GMT transfer remains fragile. These results position HumanoidArena as a benchmark for studying transferable intermediate action representations and scalable egocentric whole-body policy learning.

人形机器人分层控制全身交互仿真基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。