arXiv:2412.11484cs.AIcs.CV2024-12NeurIPS被引 18

用视觉提示集提升智能体对新环境的快速适应能力

Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents

  • 通过对比学习构建多视觉提示的注意力集成框架
  • 在多个任务上实现更高样本效率与更强泛化性能
  • 适合需要快速适应新场景的机器人与自动驾驶研究

对于与环境交互的具身强化学习智能体,快速适应未见视觉观测是理想目标,但零样本适应在强化学习中仍具挑战。本文提出一种新型对比提示集成(ConPE)框架,利用预训练视觉语言模型和一组视觉提示,实现对多种环境与物理变化的高效策略学习与适应。具体而言,设计基于引导注意力的提示集成方法,在视觉语言模型上使用多个视觉提示,构建鲁棒的状态表示;每个提示针对显著影响智能体自身视角感知的独立域因素进行对比学习。针对特定任务,注意力集成与策略联合优化,使状态表示既具备跨域泛化性,又适配任务学习。实验表明,ConPE在AI2THOR导航、egocentric-Metaworld操作及CARLA自动驾驶等多个任务上优于现有先进算法,同时提升策略学习与适应的样本效率。

原文摘要 · Abstract (English)

For embodied reinforcement learning (RL) agents interacting with the environment, it is desirable to have rapid policy adaptation to unseen visual observations, but achieving zero-shot adaptation capability is considered as a challenging problem in the RL context. To address the problem, we present a novel contrastive prompt ensemble (ConPE) framework which utilizes a pretrained vision-language model and a set of visual prompts, thus enabling efficient policy learning and adaptation upon a wide range of environmental and physical changes encountered by embodied agents. Specifically, we devise a guided-attention-based ensemble approach with multiple visual prompts on the vision-language model to construct robust state representations. Each prompt is contrastively learned in terms of an individual domain factor that significantly affects the agent's egocentric perception and observation. For a given task, the attention-based ensemble and policy are jointly learned so that the resulting state representations not only generalize to various domains but are also optimized for learning the task. Through experiments, we show that ConPE outperforms other state-of-the-art algorithms for several embodied agent tasks including navigation in AI2THOR, manipulation in egocentric-Metaworld, and autonomous driving in CARLA, while also improving the sample efficiency of policy learning and adaptation.

具身智能视觉提示强化学习快速适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。