arXiv:2604.05112cs.LGcs.AI2026-04

用预训练变压器实现可扩展的上下文强化学习,提升泛化能力。

Vintix II: Decision Pre-Trained Transformer is a Scalable In-Context Reinforcement Learner

  • 基于决策预训练的Transformer,通过流匹配进行训练
  • 在数百个任务上训练,对未见任务泛化性能显著提升
  • 适合构建通用智能体,替代专家知识蒸馏

近期在上下文强化学习(ICRL)方面的进展表明,该方法可直接在推理时训练通用智能体。算法蒸馏(AD)开创了这一范式并被扩展至多领域场景,但其对未见任务的泛化能力仍有限。决策预训练变压器(DPT)作为替代方案,在简化领域中展现出更强的上下文强化学习能力,但其可扩展性尚未验证。本文将DPT扩展至多样化的多领域环境,采用流匹配作为自然训练方式,保持其贝叶斯后验采样的解释性。结果获得一个在数百个不同任务上训练的智能体,在保留测试集上实现显著泛化提升。该智能体优于先前的AD扩展,在在线与离线推理中均表现更优,强化了ICRL作为通用智能体训练中专家蒸馏可行替代方案的地位。

原文摘要 · Abstract (English)

Recent progress in in-context reinforcement learning (ICRL) has demonstrated its potential for training generalist agents that can acquire new tasks directly at inference. Algorithm Distillation (AD) pioneered this paradigm and was subsequently scaled to multi-domain settings, although its ability to generalize to unseen tasks remained limited. The Decision Pre-Trained Transformer (DPT) was introduced as an alternative, showing stronger in-context reinforcement learning abilities in simplified domains, but its scalability had not been established. In this work, we extend DPT to diverse multi-domain environments, applying Flow Matching as a natural training choice that preserves its interpretation as Bayesian posterior sampling. As a result, we obtain an agent trained across hundreds of diverse tasks that achieves clear gains in generalization to the held-out test set. This agent improves upon prior AD scaling and demonstrates stronger performance in both online and offline inference, reinforcing ICRL as a viable alternative to expert distillation for training generalist agents.

强化学习通用智能体上下文学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。