用大模型指导强化学习探索,发现能懂目标却不会操作的瓶颈。
Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
- 让大模型看图理解任务目标,再指导智能体探索
- 在理想条件下,大模型能大幅提升早期采样效率
- 适合想用大模型做探索引导的研究者参考
强化学习中的探索问题在稀疏奖励环境下尤为困难。尽管基础模型具备强大的语义先验,但其作为零样本探索代理在经典强化学习基准上的能力尚不明确。我们在多臂老虎机、网格世界和稀疏奖励雅达利游戏上对大语言模型和视觉语言模型进行了基准测试。研究发现关键局限:虽然视觉语言模型能从视觉输入中推断高层次目标,但在精确的低层控制上始终表现不佳,即存在‘知行鸿沟’。为分析弥合此鸿沟的可能路径,我们在一个受控的理想场景中考察了一种简单的在线策略混合框架。结果表明,在该理想设置下,视觉语言模型的指导可显著提升早期阶段的样本效率,清晰揭示了利用基础模型引导探索而非端到端控制的潜力与边界。
原文摘要 · Abstract (English)
Exploration in reinforcement learning (RL) remains challenging, particularly in sparse-reward settings. While foundation models possess strong semantic priors, their capabilities as zero-shot exploration agents in classic RL benchmarks are not well understood. We benchmark LLMs and VLMs on multi-armed bandits, Gridworlds, and sparse-reward Atari to test zero-shot exploration. Our investigation reveals a key limitation: while VLMs can infer high-level objectives from visual input, they consistently fail at precise low-level control: the "knowing-doing gap". To analyze a potential bridge for this gap, we investigate a simple on-policy hybrid framework in a controlled, best-case scenario. Our results in this idealized setting show that VLM guidance can significantly improve early-stage sample efficiency, providing a clear analysis of the potential and constraints of using foundation models to guide exploration rather than for end-to-end control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。