用高效强化学习训练出更像人类的守门员,提升游戏真实感。
Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach
- 利用预收集数据和增强网络可塑性,提升训练效率。
- 守门员救球率比游戏内置AI高10%,训练速度加快50%。
- 适合资源有限的游戏工作室,已用于EA最新足球游戏。
尽管深度强化学习(DRL)在多个知名游戏平台上被用作测试平台,但该技术极少被游戏产业用于构建真实可信的AI行为。以往研究多聚焦于训练超人类水平的大型模型,这对资源有限的游戏工作室而言不切实际。本文提出一种面向工业场景的样本高效DRL方法,专门用于训练和微调游戏中的智能体。通过利用预收集数据并提升网络可塑性,该方法显著提高基于价值的DRL样本效率。我们在EA SPORTS FC 25——目前最畅销的足球模拟游戏之一中评估了该方法,训练出的守门员智能体在救球率上较游戏内置AI提升10%。消融实验表明,该方法相比标准DRL训练速度提升50%。领域专家的定性评估显示,该方法生成的行为更具人类特征,优于传统手工设计的代理。作为该方法影响力的证明,其已被应用于该系列最新版本中。
原文摘要 · Abstract (English)
While several high profile video games have served as testbeds for Deep Reinforcement Learning (DRL), this technique has rarely been employed by the game industry for crafting authentic AI behaviors. Previous research focuses on training super-human agents with large models, which is impractical for game studios with limited resources aiming for human-like agents. This paper proposes a sample-efficient DRL method tailored for training and fine-tuning agents in industrial settings such as the video game industry. Our method improves sample efficiency of value-based DRL by leveraging pre-collected data and increasing network plasticity. We evaluate our method training a goalkeeper agent in EA SPORTS FC 25, one of the best-selling football simulations today. Our agent outperforms the game's built-in AI by 10% in ball saving rate. Ablation studies show that our method trains agents 50% faster compared to standard DRL methods. Finally, qualitative evaluation from domain experts indicates that our approach creates more human-like gameplay compared to hand-crafted agents. As a testament to the impact of the approach, the method has been adopted for use in the most recent release of the series.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。