用多模态数据训练可实时响应文本指令的游戏智能体
Learning to play: A Multimodal Agent for 3D Game-Play
- 构建大规模跨游戏人类操作数据集,含文本指令
- 通过逆动力学模型推断无动作标注视频中的操作
- 在消费级显卡上实现实时推理,支持多游戏文本控制
我们认为3D第一人称视频游戏是实时多模态推理的挑战性环境。我们首先描述了一个涵盖多种3D第一人称游戏的人类游戏行为数据集,其规模和多样性均显著超过以往公开数据集,并包含文本指令。我们证明可基于该数据集学习逆动力学模型,从而在大量缺乏动作记录的公开游戏视频中推断出动作。随后,我们使用行为克隆训练了一个文本条件的游戏玩家代理,采用定制架构可在消费级GPU上实现实时推理。结果显示该模型能够应对多种3D游戏并响应文本输入。最后,我们指出若干未解决问题,如长时序任务处理及在大量游戏上的量化评估。
原文摘要 · Abstract (English)
We argue that 3-D first-person video games are a challenging environment for real-time multi-modal reasoning. We first describe our dataset of human game-play, collected across a large variety of 3-D first-person games, which is both substantially larger and more diverse compared to prior publicly disclosed datasets, and contains text instructions. We demonstrate that we can learn an inverse dynamics model from this dataset, which allows us to impute actions on a much larger dataset of publicly available videos of human game play that lack recorded actions. We then train a text-conditioned agent for game playing using behavior cloning, with a custom architecture capable of realtime inference on a consumer GPU. We show the resulting model is capable of playing a variety of 3-D games and responding to text input. Finally, we outline some of the remaining challenges such as long-horizon tasks and quantitative evaluation across a large set of games.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。