arXiv:2410.23208cs.LGcs.AI2024-10ICLR被引 45

用海量物理任务训练通用智能体,实现零样本解决新问题。

Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks

  • 通过程序生成数千万个2D物理任务,统一框架覆盖机器人控制到游戏
  • 训练的智能体在未见环境中零样本成功,微调后性能远超从零训练
  • 适合研究大规模预训练与通用强化学习的学者

尽管自监督学习在文本和图像领域已展现强大能力,但在序列决策问题中实现智能体的通用性仍面临挑战。本文提出Kinetix——一个开放式的物理强化学习环境空间,通过程序生成数千万个2D物理任务,涵盖机器人运动、抓取、视频游戏及经典RL环境。我们使用新型硬件加速物理引擎Jax2D,在训练中模拟数十亿步环境交互。训练后的智能体在2D空间中展现出强物理推理能力,可零样本解决人类设计的新环境。进一步微调该通用智能体,在目标任务上表现显著优于从头训练的RL智能体,甚至能解决标准训练完全失败的任务。结果表明,大规模混合质量在线预训练具有可行性,Kinetix为后续研究提供了有效框架。

原文摘要 · Abstract (English)

While large models trained with self-supervised learning on offline datasets have shown remarkable capabilities in text and image domains, achieving the same generalisation for agents that act in sequential decision problems remains an open challenge. In this work, we take a step towards this goal by procedurally generating tens of millions of 2D physics-based tasks and using these to train a general reinforcement learning (RL) agent for physical control. To this end, we introduce Kinetix: an open-ended space of physics-based RL environments that can represent tasks ranging from robotic locomotion and grasping to video games and classic RL environments, all within a unified framework. Kinetix makes use of our novel hardware-accelerated physics engine Jax2D that allows us to cheaply simulate billions of environment steps during training. Our trained agent exhibits strong physical reasoning capabilities in 2D space, being able to zero-shot solve unseen human-designed environments. Furthermore, fine-tuning this general agent on tasks of interest shows significantly stronger performance than training an RL agent *tabula rasa*. This includes solving some environments that standard RL training completely fails at. We believe this demonstrates the feasibility of large scale, mixed-quality pre-training for online RL and we hope that Kinetix will serve as a useful framework to investigate this further.

强化学习物理模拟通用智能体预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。