用分而治之策略让大模型高效决策,解决长期任务难题
Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning
- 高层策略生成抽象计划,低层控制器按步骤执行
- 在ScienceWorld和ALFWorld上显著提升长期任务表现
- 适合需要快速适应变化环境的智能体应用
尽管具备复杂推理能力,大型语言模型(LLMs)在长时序决策任务中仍因探索不足和长期信用分配困难而表现不佳,尤其在稀疏奖励场景下。受分而治之思想启发,我们提出新型框架GLIDER(Grounding Language Models as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning),为LLM策略引入一种参数高效且通用的层次结构。低层控制器通过高层策略生成的抽象、逐步计划进行监督学习,将复杂问题分解为一系列连贯的思维链子任务,提供灵活的时间抽象,显著增强长时序任务的探索与学习效果。此外,得益于任务无关的低层技能强迁移性,GLIDER支持快速在线适应非平稳环境。在ScienceWorld和ALFWorld基准上的实验表明,GLIDER实现了稳定的性能提升,并增强了泛化能力。
原文摘要 · Abstract (English)
While showing sophisticated reasoning abilities, large language models (LLMs) still struggle with long-horizon decision-making tasks due to deficient exploration and long-term credit assignment, especially in sparse-reward scenarios. Inspired by the divide-and-conquer principle, we propose an innovative framework **GLIDER** (**G**rounding **L**anguage Models as Eff**I**cient **D**ecision-Making Agents via Offline Hi**E**rarchical **R**einforcement Learning) that introduces a parameter-efficient and generally applicable hierarchy to LLM policies. We develop a scheme where the low-level controller is supervised with abstract, step-by-step plans that are learned and instructed by the high-level policy. This design decomposes complicated problems into a series of coherent chain-of-thought reasoning sub-tasks, providing flexible temporal abstraction to significantly enhance exploration and learning for long-horizon tasks. Furthermore, GLIDER facilitates fast online adaptation to non-stationary environments owing to the strong transferability of its task-agnostic low-level skills. Experiments on ScienceWorld and ALFWorld benchmarks show that GLIDER achieves consistent performance gains, along with enhanced generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。