arXiv:2510.23691cs.AI2025-10被引 17

用键盘鼠标统一动作空间,训练可通用的游戏智能体

Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents

  • 以人类输入为锚点构建统一动作空间,支持跨游戏持续预训练
  • 在开放世界Minecraft中成功率是之前模型的2倍,接近真人水平
  • 适合想构建通用计算代理的研究者与开发者

我们提出Game-TARS,一种基于统一、可扩展动作空间的通用游戏智能体,该空间以人类对齐的原生键盘鼠标输入为锚点。不同于依赖API或GUI的方法,此范式支持在操作系统、网页和仿真游戏等异构领域进行大规模持续预训练。Game-TARS在超过5000亿个标记上进行预训练,涵盖多样化轨迹与多模态数据。关键技术包括降低因果混淆的衰减持续损失,以及平衡推理深度与推理成本的高效稀疏思维策略。实验表明,Game-TARS在开放世界Minecraft任务中的成功率约为先前SOTA模型的2倍,在未见过的网页3D游戏中表现接近未经训练的人类,且在FPS基准测试中优于GPT-5、Gemini-2.5-Pro和Claude-4-Sonnet。训练时长与测试时长的扩展结果证实,统一动作空间在跨游戏和多模态数据扩展下仍能持续提升性能。结果表明,简单、可扩展的动作表示结合大规模预训练,为具备广泛计算机使用能力的通用智能体提供了有前景的路径。

原文摘要 · Abstract (English)

We present Game-TARS, a generalist game agent trained with a unified, scalable action space anchored to human-aligned native keyboard-mouse inputs. Unlike API- or GUI-based approaches, this paradigm enables large-scale continual pre-training across heterogeneous domains, including OS, web, and simulation games. Game-TARS is pre-trained on over 500B tokens with diverse trajectories and multimodal data. Key techniques include a decaying continual loss to reduce causal confusion and an efficient Sparse-Thinking strategy that balances reasoning depth and inference cost. Experiments show that Game-TARS achieves about 2 times the success rate over the previous sota model on open-world Minecraft tasks, is close to the generality of fresh humans in unseen web 3d games, and outperforms GPT-5, Gemini-2.5-Pro, and Claude-4-Sonnet in FPS benchmarks. Scaling results on training-time and test-time confirm that the unified action space sustains improvements when scaled to cross-game and multimodal data. Our results demonstrate that simple, scalable action representations combined with large-scale pre-training provide a promising path toward generalist agents with broad computer-use abilities.

通用智能体游戏代理多模态预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。