arXiv:2511.19584cs.LGcs.CV2025-11被引 12

一个模型搞定200个任务,实现高效多任务连续控制。

Learning Massively Multitask World Models for Continuous Control

  • 用演示数据预训练+在线交互联合优化,构建多任务世界模型。
  • 在200个任务上表现优于基线,适应新任务速度快、数据效率高。
  • 适合研究通用智能体与多任务强化学习的学者和开发者。

通用控制需要能在多种任务和身体形态上运作的智能体,但当前连续控制的强化学习研究仍以单任务或离线训练为主,似乎暗示在线强化学习难以扩展。受基础模型范式(大规模预训练后轻量微调)启发,我们探究是否可让单一智能体通过在线交互学习数百个任务。为此,我们提出一个新基准,包含200个跨领域、跨形态的多样化任务,每个任务配有语言指令、示范数据,部分含图像观测。我们进一步提出 extit{Newt}——一种语言条件的多任务世界模型:先在示范数据上预训练,获得任务感知表征与动作先验,再在所有任务上联合进行在线交互优化。实验表明,Newt 在多任务性能与数据效率方面优于多个强基线,表现出优异的开环控制能力,并能快速适应未见任务。我们开源了环境、示范数据、训练与评估代码及200多个检查点。

原文摘要 · Abstract (English)

General-purpose control demands agents that act across many tasks and embodiments, yet research on reinforcement learning (RL) for continuous control remains dominated by single-task or offline regimes, reinforcing a view that online RL does not scale. Inspired by the foundation model recipe (large-scale pretraining followed by light RL) we ask whether a single agent can be trained on hundreds of tasks with online interaction. To accelerate research in this direction, we introduce a new benchmark with 200 diverse tasks spanning many domains and embodiments, each with language instructions, demonstrations, and optionally image observations. We then present \emph{Newt}, a language-conditioned multitask world model that is first pretrained on demonstrations to acquire task-aware representations and action priors, and then jointly optimized with online interaction across all tasks. Experiments show that Newt yields better multitask performance and data-efficiency than a set of strong baselines, exhibits strong open-loop control, and enables rapid adaptation to unseen tasks. We release our environments, demonstrations, code for training and evaluation, as well as 200+ checkpoints.

多任务学习强化学习世界模型连续控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。