基于4万小时游戏视频训练的通用游戏智能体,能跨游戏高效迁移。
NitroGen: An Open Foundation Model for Generalist Gaming Agents
- 用互联网规模的游戏视频自动提取操作数据,构建多游戏训练集。
- 在未见过游戏中任务成功率提升最高达52%,显著优于从零训练模型。
- 适合研究通用智能体、游戏AI及具身智能的开发者和研究人员。
我们提出NitroGen,一个面向通用游戏智能体的视觉-动作基础模型,其在超过1000款游戏的4万小时游戏视频上进行训练。核心包括:1)通过自动解析公开游戏视频提取玩家操作,构建互联网规模的视频-动作数据集;2)设计支持跨游戏泛化评估的多游戏基准环境;3)采用大规模行为克隆训练统一的视觉-动作模型。NitroGen在多种场景中表现出色,涵盖3D动作类游戏中的战斗应对、2D平台游戏的高精度控制以及程序生成世界的探索。在未见过的游戏上具备强迁移能力,任务成功率相对从零训练的模型最高提升52%。我们开放数据集、评估套件及模型权重,推动通用具身智能体的研究进展。
原文摘要 · Abstract (English)
We introduce NitroGen, a vision-action foundation model for generalist gaming agents that is trained on 40,000 hours of gameplay videos across more than 1,000 games. We incorporate three key ingredients: 1) an internet-scale video-action dataset constructed by automatically extracting player actions from publicly available gameplay videos, 2) a multi-game benchmark environment that can measure cross-game generalization, and 3) a unified vision-action model trained with large-scale behavior cloning. NitroGen exhibits strong competence across diverse domains, including combat encounters in 3D action games, high-precision control in 2D platformers, and exploration in procedurally generated worlds. It transfers effectively to unseen games, achieving up to 52% relative improvement in task success rates over models trained from scratch. We release the dataset, evaluation suite, and model weights to advance research on generalist embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。