用视觉语言模型指导探索,让智能体学会有意义的高阶行为。
SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models
- 从视觉语言模型中提取有趣性奖励信号,驱动智能体探索
- 在机器人和游戏模拟中仅用图像和底层动作就发现多样行为
- 无需语言环境或高层动作,适合通用世界模型学习
探索是强化学习的核心。内在动机试图将探索与外部任务奖励解耦,但现有方法多聚焦低层次交互。儿童玩耍表明,他们通过模仿或与照料者互动来开展有意义的高层次行为。近期工作尝试利用基础模型注入语义偏见,但常依赖不现实假设(如嵌入语言的环境或高层动作)。本文提出 SENSEI 框架,使基于模型的强化学习智能体具备对语义上合理行为的内在动机。SENSEI 从视觉语言模型(VLM)注释中提炼有趣性奖励信号,使智能体通过世界模型预测该信号。基于模型的强化学习训练探索策略,同时最大化语义奖励与不确定性。在机器人及类游戏仿真环境中,SENSEI 仅使用图像观测与低层动作,便发现了多种有意义的行为。该方法为从基础模型反馈中学习提供了通用工具,具有重要研究价值。
原文摘要 · Abstract (English)
Exploration is a cornerstone of reinforcement learning (RL). Intrinsic motivation attempts to decouple exploration from external, task-based rewards. However, established approaches to intrinsic motivation that follow general principles such as information gain, often only uncover low-level interactions. In contrast, children's play suggests that they engage in meaningful high-level behavior by imitating or interacting with their caregivers. Recent work has focused on using foundation models to inject these semantic biases into exploration. However, these methods often rely on unrealistic assumptions, such as language-embedded environments or access to high-level actions. We propose SEmaNtically Sensible ExploratIon (SENSEI), a framework to equip model-based RL agents with an intrinsic motivation for semantically meaningful behavior. SENSEI distills a reward signal of interestingness from Vision Language Model (VLM) annotations, enabling an agent to predict these rewards through a world model. Using model-based RL, SENSEI trains an exploration policy that jointly maximizes semantic rewards and uncertainty. We show that in both robotic and video game-like simulations SENSEI discovers a variety of meaningful behaviors from image observations and low-level actions. SENSEI provides a general tool for learning from foundation model feedback, a crucial research direction, as VLMs become more powerful.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。