大模型无需微调即可高效建模奖励,提升决策能力。
On the Modeling Capabilities of Large Language Models for Sequential Decision Making
- 用AI反馈生成奖励信号,间接指导强化学习。
- 无需微调,大模型在复杂任务中表现优异,提升探索与信用分配。
- 合成数据微调可增强对陌生环境的适应性,避免遗忘。
大型预训练模型在多模态推理与规划任务中表现日益出色,为复杂序列决策问题提供了新可能。本文研究了大语言模型(LLMs)在多种交互式领域中用于强化学习(RL)的能力。我们评估了其直接生成动作或先生成奖励模型再训练智能体的策略。结果表明,即使不进行特定任务微调,大模型在奖励建模方面也表现出色。尤其通过人工智能反馈构建奖励,是最具通用性的方法,能显著改善信用分配与探索效率。在动态未知环境中,利用合成数据对大模型进行微调,可显著提升其奖励建模能力,同时缓解灾难性遗忘,进一步拓展其在序列决策任务中的应用潜力。
原文摘要 · Abstract (English)
Large pretrained models are showing increasingly better performance in reasoning and planning tasks across different modalities, opening the possibility to leverage them for complex sequential decision making problems. In this paper, we investigate the capabilities of Large Language Models (LLMs) for reinforcement learning (RL) across a diversity of interactive domains. We evaluate their ability to produce decision-making policies, either directly, by generating actions, or indirectly, by first generating reward models to train an agent with RL. Our results show that, even without task-specific fine-tuning, LLMs excel at reward modeling. In particular, crafting rewards through artificial intelligence (AI) feedback yields the most generally applicable approach and can enhance performance by improving credit assignment and exploration. Finally, in environments with unfamiliar dynamics, we explore how fine-tuning LLMs with synthetic data can significantly improve their reward modeling capabilities while mitigating catastrophic forgetting, further broadening their utility in sequential decision-making tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。