arXiv:2410.03303cs.LGcs.CV2024-10被引 8

让智能体在未知环境里自我学习,提升理解与决策能力。

SELU: Self-Learning Embodied MLLMs in Unknown Environments

  • 设计自问自答与回溯重标注机制,让模型从自身经验中提炼知识。
  • 在AI2-THOR和VirtualHome中,模型环境理解力提升28%~30%,决策力提升20%~24%。
  • 适合研究自主智能体、自学习系统或强化学习中的多模态应用者。

近期,多模态大语言模型(MLLMs)展现出强大的视觉理解与决策能力,推动了在未知环境中自主优化MLLMs的探索。然而,外部反馈(如人类或环境反馈)并不总能获取。现有方法主要通过投票与评分机制增强MLLMs的决策能力,却较少关注其在未知环境中的环境理解能力提升。为充分释放MLLMs的自学习潜力,我们提出一种受强化学习中演员-评论家范式启发的新方法——SELU。评论家通过自问与回溯重标注,从演员收集的交互轨迹中提取知识,从而增强环境理解;同时,演员通过评论家提供的自反馈得到改进,提升决策能力。我们在AI2-THOR和VirtualHome环境中评估该方法,结果显示,通过自学习,评论家性能提升约28%和30%,演员性能提升约20%和24%。

原文摘要 · Abstract (English)

Recently, multimodal large language models (MLLMs) have demonstrated strong visual understanding and decision-making capabilities, enabling the exploration of autonomously improving MLLMs in unknown environments. However, external feedback like human or environmental feedback is not always available. To address this challenge, existing methods primarily focus on enhancing the decision-making capabilities of MLLMs through voting and scoring mechanisms, while little effort has been paid to improving the environmental comprehension of MLLMs in unknown environments. To fully unleash the self-learning potential of MLLMs, we propose a novel actor-critic self-learning paradigm, dubbed SELU, inspired by the actor-critic paradigm in reinforcement learning. The critic employs self-asking and hindsight relabeling to extract knowledge from interaction trajectories collected by the actor, thereby augmenting its environmental comprehension. Simultaneously, the actor is improved by the self-feedback provided by the critic, enhancing its decision-making. We evaluate our method in the AI2-THOR and VirtualHome environments, and SELU achieves critic improvements of approximately 28% and 30%, and actor improvements of about 20% and 24% via self-learning.

自学习多模态强化学习智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。