arXiv:2607.15163cs.ROcs.AI2026-07被引 1

通过协调学习范式、数据多样性和模型架构,提升人形机器人控制性能。

Scaling Behavior Foundation Model for Humanoid Robots

论文配图:Scaling Behavior Foundation Model for Humanoid Robots
图 1 · 摘自论文原文
  • 将多种控制问题统一为全局坐标下的全身行为复现。
  • 实测在全局模式下关键点误差降低82%,局部模式降低10%。
  • 适合希望构建通用、可扩展人形机器人控制系统的研究者。

人形机器人控制需要自然的全身协调、对控制信号的精确实时响应以及在不同环境中的鲁棒泛化能力,是通用具身智能体的核心挑战。行为基础模型(BFMs)近期成为解决这一难题的有前景方案,通过利用大规模行为数据实现更强的表现力、多样性和泛化能力。然而,尽管对扩展BFMs的兴趣日益增长,如何协调学习范式、行为数据与模型架构以实现有效扩展仍不明确。本文重新审视了BFMs的扩展策略,证明通过三个核心组件的协同可显著提升性能:1)运动追踪学习范式,将多样化的控制问题重构为全局坐标系中整体行为的再现;2)在线策略回放数量与参考动作多样性的战略协同;3)名为人形变压器(Humanoid Transformer)的表达性强且可扩展的模型架构,促进结构化行为表示的自然涌现。在仿真和真实世界部署中广泛实验表明,该方法显著提升了控制保真度和任务泛化能力,相比现有控制器,测试集上的平均每关键点位置误差(MPKPE)在本地模式下降低超10%,全局模式下降低82%。这些结果确立了BFM作为可扩展、通用的人形机器人控制原则性有效基础。

原文摘要 · Abstract (English)

Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to further improve their capabilities, it remains unclear how key factors, including the learning paradigm, behavioral data and model architecture should be coordinated to enable effective scaling. In this work, we revisit the scaling recipe for BFMs and demonstrate that substantial performance gains can be achieved through the coordination of three core components: 1) the learning paradigm of motion tracking that reformulates diverse humanoid control problems as the reproduction of integrated whole-body behaviors in the global frame; 2) the strategic synergy between on-policy rollout quantity and reference motion diversity; and 3) the expressive and scalable model architecture termed Humanoid Transformer that facilitates the natural emergence of structured behavioral representations. Through extensive experiments in both simulation and real-world deployment, we demonstrate that our approach yields significant improvements in control fidelity and task generalization, reducing Mean Per-Keypoint Position Error (MPKPE) on the test set by over 10% in local mode and 82% in global mode compared with existing humanoid controllers. These results establish BFM as a principled and effective foundation for scalable and general-purpose humanoid control.

人形机器人行为建模可扩展性控制优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。