用双尺度世界模型提升大模型在复杂探索任务中的学习效率。
Dual-Scale World Models for LLM Agents Towards Hard-Exploration Problems
- 全局维护高价值发现轨迹,局部通过多路径优势反思引导试错。
- 在文本游戏基准上达到新最优,环境交互量仅为强化学习方法的1/100至1/800。
- 适合需要深度探索与高效学习的复杂决策场景,如智能体自主探索。
基于大模型的智能体虽有进展,但在需通过探索获取新知识的“硬探索”任务中仍受限。本文提出GLoW,一种利用双尺度世界模型的新方法:在全局尺度维持高价值发现轨迹,在局部尺度通过多路径优势反射机制,从试错中推断基于优势的进展信号以指导探索。为评估该框架在硬探索任务中的表现,我们在文本游戏基准杰里科(Jericho)上进行测试,结果表明,GLoW在基于大模型的方法中达到新最优性能;相比最先进强化学习方法,其性能相当,但环境交互次数减少100至800倍。
原文摘要 · Abstract (English)
LLM-based agents have seen promising advances, yet they are still limited in "hard-exploration" tasks requiring learning new knowledge through exploration. We present GLoW, a novel approach leveraging dual-scale world models, maintaining a trajectory frontier of high-value discoveries at the global scale, while learning from local trial-and-error in exploration through a Multi-path Advantage Reflection mechanism which infers advantage-based progress signals to guide exploration. To evaluate our framework for hard-exploration, we tackle the Jericho benchmark suite of text-based games, where GLoW achieves a new state-of-theart performance for LLM-based approaches. Compared to state-of-the-art RLbased methods, our approach achieves comparable performance while requiring 100-800x fewer environment interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。