arXiv:2505.11221cs.LG2025-05被引 3

用大模型指导强化学习,让智能体少试多学。

Sample Efficient Reinforcement Learning via Large Vision Language Model Distillation

  • 用大视觉语言模型当老师,指导强化学习减少无效探索。
  • 实验显示,新方法使基线算法样本效率大幅提升。
  • 无需手动描述环境,适合多种复杂任务快速部署。

近期研究显示,多模态基础模型在应对复杂决策问题上具有潜力。然而,其庞大的参数量导致实际部署资源消耗大,对资源受限系统不友好。强化学习(RL)虽有潜力构建特定任务智能体,但存在高样本复杂性,限制了实际应用。为此,我们提出LVLM2P框架,将大视觉语言模型(LVLM)的知识蒸馏到更高效的强化学习智能体中。该方法利用LVLM作为教师,基于强化学习智能体收集的轨迹提供指导性动作,有效减少学习初期无意义的探索,显著加速学习进程。同时,通过直接从视觉观察中建议动作,无需人工编写环境文本描述,提升了在多样化任务中的适用性。实验表明,LVLM2P显著提升了基线强化学习算法的样本效率。

原文摘要 · Abstract (English)

Recent research highlights the potential of multimodal foundation models in tackling complex decision-making challenges. However, their large parameters make real-world deployment resource-intensive and often impractical for constrained systems. Reinforcement learning (RL) shows promise for task-specific agents but suffers from high sample complexity, limiting practical applications. To address these challenges, we introduce LVLM to Policy (LVLM2P), a novel framework that distills knowledge from large vision-language models (LVLM) into more efficient RL agents. Our approach leverages the LVLM as a teacher, providing instructional actions based on trajectories collected by the RL agent, which helps reduce less meaningful exploration in the early stages of learning, thereby significantly accelerating the agent's learning progress. Additionally, by leveraging the LVLM to suggest actions directly from visual observations, we eliminate the need for manual textual descriptors of the environment, enhancing applicability across diverse tasks. Experiments show that LVLM2P significantly enhances the sample efficiency of baseline RL algorithms.

强化学习大模型蒸馏样本高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。