LLM能在上下文里通过外部奖励在线学习,像智能体一样决策。
LLMs Are In-Context Bandit Reinforcement Learners
- 用外部奖励替代标注数据,在上下文中实时学习
- 从500M到700亿参数模型均展现学习能力
- 适合研究大模型在线决策与错误推理局限的学者
大型语言模型(LLMs)在上下文学习(ICL)中表现优异,这是一种依赖上下文添加标注样例的监督学习方法。本文研究了上下文强化学习(ICRL)的一种情境化带状问题形式,即模型在上下文中、在线地从外部奖励中学习,而非依赖监督数据。实验表明,LLMs能有效实现此类学习,并在多个具有挑战性的分类任务上,对从500M到70B参数规模的模型进行了详细分析。研究识别并缓解了该过程中的不稳定性,验证了模型可使用语义和抽象标签进行学习,并揭示了其学习的缩放趋势。结果凸显了LLMs在上下文强化学习方面的能力,同时也指出了其在隐式错误推理方面的根本局限。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel at in-context learning (ICL), a supervised learning technique that relies on adding annotated examples to the model context. We investigate a contextual bandit version of in-context reinforcement learning (ICRL), where models learn in-context, online, from external reward, instead of supervised data. We show that LLMs effectively demonstrate such learning, and provide a detailed study of the phenomena, experimenting with challenging classification tasks and models of sizes from 500M to 70B parameters. This includes identifying and addressing the instability of the process, demonstrating learning with both semantic and abstract labels, and showing scaling trends. Our findings highlight ICRL capabilities in LLMs, while also underscoring fundamental limitations in their implicit reasoning about errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。