arXiv:2412.08442cs.LG2024-12CVPR被引 49

将多模态大模型改造为通用具身智能体,实现跨领域任务泛化。

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons

  • 用多具身动作标记器统一建模不同场景下的智能体行为。
  • 在仿真环境中通过在线强化学习训练,实现跨任务泛化性能超越基准模型。
  • 适合研究通用人工智能、具身智能与多模态交互的开发者参考。

我们考察多模态大语言模型(MLLM)在传统语言与视觉任务之外的多样化领域中的能力,重点关注具身智能、游戏、用户界面控制与规划。为此,我们提出将MLLM适配为通用具身智能体(GEA)。GEA是一个单一统一模型,通过多具身动作标记器实现跨领域的任务定位。该模型在大规模具身经验数据集上采用监督学习,并在交互式模拟器中进行在线强化学习(RL)训练。我们探讨了构建此类模型所需的数据与算法选择。研究发现,跨领域数据与在线强化学习对构建通用智能体至关重要。最终的GEA模型在多个未见过的任务上表现出强泛化能力,优于其他通用模型及特定任务基准方法。

原文摘要 · Abstract (English)

We examine the capability of Multimodal Large Language Models (MLLMs) to tackle diverse domains that extend beyond the traditional language and vision tasks these models are typically trained on. Specifically, our focus lies in areas such as Embodied AI, Games, UI Control, and Planning. To this end, we introduce a process of adapting an MLLM to a Generalist Embodied Agent (GEA). GEA is a single unified model capable of grounding itself across these varied domains through a multi-embodiment action tokenizer. GEA is trained with supervised learning on a large dataset of embodied experiences and with online RL in interactive simulators. We explore the data and algorithmic choices necessary to develop such a model. Our findings reveal the importance of training with cross-domain data and online RL for building generalist agents. The final GEA model achieves strong generalization performance to unseen tasks across diverse benchmarks compared to other generalist models and benchmark-specific approaches.

具身智能多模态模型强化学习通用智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。