用视觉语言模型生成概念,引导强化学习高效探索目标物体。
CDE: Concept-Driven Exploration for Reinforcement Learning
- 基于文本描述生成物体概念,作为探索的弱监督信号。
- 通过概念重建误差获得内在奖励,提升探索效率。
- 在仿真与真实机械臂上均表现稳定,适合复杂视觉任务。
智能探索仍是强化学习中的关键挑战,尤其在视觉控制任务中。与低维状态强化学习不同,视觉强化学习需从原始像素中提取任务相关结构,导致探索效率低下。本文提出概念驱动探索(CDE),利用预训练视觉语言模型(VLM)根据文本任务描述生成以物体为中心的视觉概念,作为弱但可能含噪的监督信号。CDE不直接依赖这些噪声信号,而是训练策略通过辅助目标重构概念,学习概念的通用表示,并将重构精度作为内在奖励,引导探索聚焦于任务相关物体。在五个具有挑战性的模拟视觉操作任务中,CDE实现了高效且精准的探索,并对合成错误和噪声的VLM预测保持鲁棒性。最后,我们在Franka机械臂上实现了真实世界迁移,成功完成一项实际操作任务,成功率高达80%。
原文摘要 · Abstract (English)
Intelligent exploration remains a critical challenge in reinforcement learning (RL), especially in visual control tasks. Unlike low-dimensional state-based RL, visual RL must extract task-relevant structure from raw pixels, making exploration inefficient. We propose Concept-Driven Exploration (CDE), which leverages a pre-trained vision-language model (VLM) to generate object-centric visual concepts from textual task descriptions as weak, potentially noisy supervisory signals. Rather than directly conditioning on these noisy signals, CDE trains a policy to reconstruct the concepts via an auxiliary objective, learning general representations of the concepts and using reconstruction accuracy as an intrinsic reward to guide exploration toward task-relevant objects. Across five challenging simulated visual manipulation tasks, CDE achieves efficient, targeted exploration and remains robust to both synthetic errors and noisy VLM predictions. Finally, we demonstrate real-world transfer by deploying CDE on a Franka arm, attaining an 80\% success rate in a real-world manipulation task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。