arXiv:2604.07993cs.RO2026-04被引 9

让机器人像人一样用全身协调完成复杂操作

HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

论文配图:HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation
图 1 · 摘自论文原文
  • 用统一状态表示实现不同机器人的可扩展学习
  • 在真实机器人上达成最高成功率和泛化能力
  • 适合做高自由度人形机器人控制的研究者

人类通过全身协同完成复杂操作,而大多数视觉-语言-动作(VLA)模型将机器人肢体独立处理,导致高自由度人形机器人控制困难且不稳定。我们提出HEX,一种面向全尺寸双足人形机器人的状态中心框架。HEX引入一种与人形对齐的通用状态表示,实现跨异构机体的可扩展学习;并采用混合专家统一本体感知预测器,从大规模多机体轨迹数据中建模全身协调与时间动态。为高效捕捉时序视觉上下文,HEX使用轻量级历史标记总结过往观测,避免推理中重复编码历史图像。同时结合残差门控融合机制与流匹配动作头,自适应融合视觉-语言线索与本体动态以生成动作。真实世界人形操作任务实验表明,HEX在任务成功率和泛化能力上达到当前最优,尤其在快速反应与长时程场景中表现突出。

原文摘要 · Abstract (English)

Humans achieve complex manipulation through coordinated whole-body control, whereas most Vision-Language-Action (VLA) models treat robot body parts largely independently, making high-DoF humanoid control challenging and often unstable. We present HEX, a state-centric framework for coordinated manipulation on full-sized bipedal humanoid robots. HEX introduces a humanoid-aligned universal state representation for scalable learning across heterogeneous embodiments, and incorporates a Mixture-of-Experts Unified Proprioceptive Predictor to model whole-body coordination and temporal motion dynamics from large-scale multi-embodiment trajectory data. To efficiently capture temporal visual context, HEX uses lightweight history tokens to summarize past observations, avoiding repeated encoding of historical images during inference. It further employs a residual-gated fusion mechanism with a flow-matching action head to adaptively integrate visual-language cues with proprioceptive dynamics for action generation. Experiments on real-world humanoid manipulation tasks show that HEX achieves state-of-the-art performance in task success rate and generalization, particularly in fast-reaction and long-horizon scenarios.

人形机器人全身控制多智能体动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。