arXiv:2603.01452cs.AIcs.RO2026-03

通过多任务学习提升人形机器人控制能力,更高效地实现通用技能掌握。

Scaling Tasks, Not Samples: Mastering Humanoid Control through Multi-Task Model-Based Reinforcement Learning

  • 用共享世界模型整合多任务经验,提升学习效率
  • 在HumanoidBench上以更少样本达到顶尖性能
  • 适合追求高效在线学习的机器人研究者

开发能掌握多种技能的通用机器人仍是具身智能的核心挑战。尽管近期进展聚焦于扩大模型参数和离线数据集,但这类方法在机器人领域受限于需主动交互的学习需求。我们提出,有效的在线学习应侧重于扩大任务数量,而非单个任务的数据量。该范式凸显了基于模型强化学习(MBRL)的结构优势:物理动力学在不同任务间具有不变性,共享的世界模型可聚合多任务经验,学习出鲁棒的、任务无关的表征。相比之下,无模型方法在相似状态下需执行冲突动作时会遭遇梯度干扰。因此,任务多样性对MBRL起到了正则化作用,提升动力学建模能力和样本效率。我们提出高效零多任务算法(EZ-M),一种面向在线学习的样本高效多任务MBRL方法。在挑战性的全身控制基准HumanoidBench上,EZ-M展现出当前最优性能,且样本效率显著优于强基线,无需极端参数扩展。这些结果确立了任务扩展作为可扩展机器人学习的关键维度。

原文摘要 · Abstract (English)

Developing generalist robots capable of mastering diverse skills remains a central challenge in embodied AI. While recent progress emphasizes scaling model parameters and offline datasets, such approaches are limited in robotics, where learning requires active interaction. We argue that effective online learning should scale the \emph{number of tasks}, rather than the number of samples per task. This regime reveals a structural advantage of model-based reinforcement learning (MBRL). Because physical dynamics are invariant across tasks, a shared world model can aggregate multi-task experience to learn robust, task-agnostic representations. In contrast, model-free methods suffer from gradient interference when tasks demand conflicting actions in similar states. Task diversity therefore acts as a regularizer for MBRL, improving dynamics learning and sample efficiency. We instantiate this idea with \textbf{EfficientZero-Multitask (EZ-M)}, a sample-efficient multi-task MBRL algorithm for online learning. Evaluated on \textbf{HumanoidBench}, a challenging whole-body control benchmark, EZ-M achieves state-of-the-art performance with significantly higher sample efficiency than strong baselines, without extreme parameter scaling. These results establish task scaling as a critical axis for scalable robotic learning. The project website is available \href{https://yewr.github.io/ez_m/}{here}.

机器人控制多任务学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。