大模型行为克隆提升因果推理,实现实时游戏智能
Scaling Behavior Cloning Improves Causal Reasoning: An Open Model for Real-Time Video Game Playing
- 用大规模数据和模型训练游戏行为克隆模型
- 12亿参数模型在3D游戏中表现媲美人类玩家
- 模型越大越能学会因果推理,适合实时游戏应用
行为克隆因模型与数据规模扩大而重新兴起。本文提出一个开放训练方案,用于构建可在消费级显卡上实时推理的视频游戏基础模型。我们公开超过8300小时高质量人类游戏数据、训练与推理代码及预训练权重。实验表明,最优模型在多种3D游戏中表现接近人类水平。通过该方案研究行为克隆的缩放规律,聚焦因果推理。在受控玩具环境中,我们证明增加训练数据和网络深度可促使模型学习更优因果策略。在真实规模下验证,分析了高达12亿参数的模型,发现玩具环境中的因果提升随模型规模与训练步数增长依然有效。
原文摘要 · Abstract (English)
Behavior cloning has seen a resurgence as scaling model and data sizes demonstrate strong performance. In this work, we introduce an open recipe for training a video game playing foundation model designed for inference in realtime on a consumer GPU. We release all data (8300+ hours of high quality human gameplay), training and inference code, and pretrained checkpoints under an open license. Empirically, we show that our best model achieves performance competitive with human players across a variety of 3D games. We use this recipe to investigate the scaling laws of behavior cloning, with a focus on causal reasoning. In a controlled toy setting, we first demonstrate that increasing training data and network depth leads to the model learning a more causal policy. We then validate these findings at scale, analyzing models up to 1.2 billion parameters. We observe that the causal improvements seen in the toy domain hold true as model size and training steps increase.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。