arXiv:2501.19400cs.LGcs.AI2025-01ICML被引 12

用上下文强化学习让模型实时学会新动作,跨领域通用性强。

Vintix: Action Model via In-Context Reinforcement Learning

  • 通过算法蒸馏构建可在上下文中自适应学习的动作模型。
  • 在多领域任务中表现优于传统专家蒸馏方法。
  • 适合需要实时决策与跨场景适应的通用智能系统研究者。

上下文强化学习(ICRL)是一种有前景的范式,使通用智能体在推理时通过试错交互学习行为,类似大语言模型的上下文适应,但聚焦于奖励最大化。然而,ICRL在超越简单任务和单领域设置下的可扩展性仍是开放挑战。本文首次推进了ICRL的规模化,提出一个固定、跨领域的动作模型,可通过上下文强化学习习得行为。结果表明,为促进ICRL而设计的算法蒸馏框架,是构建多功能动作模型的有力且具有竞争力的替代方案。这些发现凸显了ICRL作为可扩展通用决策系统方法的潜力。代码已发布于 https://github.com/dunnolab/vintix。

原文摘要 · Abstract (English)

In-Context Reinforcement Learning (ICRL) represents a promising paradigm for developing generalist agents that learn at inference time through trial-and-error interactions, analogous to how large language models adapt contextually, but with a focus on reward maximization. However, the scalability of ICRL beyond toy tasks and single-domain settings remains an open challenge. In this work, we present the first steps toward scaling ICRL by introducing a fixed, cross-domain model capable of learning behaviors through in-context reinforcement learning. Our results demonstrate that Algorithm Distillation, a framework designed to facilitate ICRL, offers a compelling and competitive alternative to expert distillation to construct versatile action models. These findings highlight the potential of ICRL as a scalable approach for generalist decision-making systems. Code released at https://github.com/dunnolab/vintix

强化学习通用智能上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。