arXiv:2502.15214cs.LGcs.AI2025-02IJCAI被引 14

将大模型与视觉语言模型融入强化学习,解决决策难题。

The Evolving Landscape of LLM- and VLM-Integrated Reinforcement Learning

  • 用大模型充当智能体、规划者或奖励函数
  • 提升长周期任务规划与奖励设计能力
  • 适合研究多模态强化学习的学者参考

强化学习在序列决策任务中表现优异。与此同时,大型语言模型(LLMs)和视觉语言模型(VLMs)在多模态理解与推理方面展现出强大能力。这些进展推动了将 LLMs 与 VLMs 融入强化学习的研究热潮。本文综述了代表性工作,展示如何利用大模型克服强化学习中的关键挑战:先验知识不足、长时程规划困难以及奖励设计复杂。我们提出一个分类框架,将大模型/视觉语言模型辅助的强化学习方法分为三种角色:智能体、规划者和奖励函数。最后探讨了开放问题,包括模型对齐性、偏见缓解、表征优化与动作建议。通过整合现有研究并指明未来方向,本文建立了一个融合自然语言与视觉理解的强化学习框架,推动多模态决策系统的发展。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has shown impressive results in sequential decision-making tasks. Meanwhile, Large Language Models (LLMs) and Vision-Language Models (VLMs) have emerged, exhibiting impressive capabilities in multimodal understanding and reasoning. These advances have led to a surge of research integrating LLMs and VLMs into RL. In this survey, we review representative works in which LLMs and VLMs are used to overcome key challenges in RL, such as lack of prior knowledge, long-horizon planning, and reward design. We present a taxonomy that categorizes these LLM/VLM-assisted RL approaches into three roles: agent, planner, and reward. We conclude by exploring open problems, including grounding, bias mitigation, improved representations, and action advice. By consolidating existing research and identifying future directions, this survey establishes a framework for integrating LLMs and VLMs into RL, advancing approaches that unify natural language and visual understanding with sequential decision-making.

强化学习大模型多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。