arXiv:2607.00642cs.AIcs.LG2026-07

让游戏AI实时按风格操控,玩家可自由调整行为表现。

Coachable agents for interactive gameplay

论文配图:Coachable agents for interactive gameplay
图 1 · 摘自论文原文
  • 用通用价值函数与训练场景设计,实现风格化强化学习
  • 在三类不同游戏中均保持风格一致性与任务完成度
  • 支持运行时动态选择行为风格,适合交互式应用

强化学习在游戏、机器人和基础模型等领域已展现强大能力,通常通过试错学习单一最优策略。但在许多场景中,用户希望对任务执行方式施加实时控制,我们称这种控制为“风格”。本文结合通用价值函数近似器(UVFAs)与精心设计的训练场景、学习算法及数据增强,构建了一个可教练的游戏智能体框架。该框架在《地平线:零之曙光》《GT赛车》及一个开源人形机器人测试环境中的应用表明,尽管领域差异大——从赛车竞速、风格化战斗到人形行走——各智能体均能保持对风格指令的高度一致性,并顺利完成核心任务。关键在于,该方法允许终端用户在运行时灵活选择最终行为,实现对智能体表现的动态控制。

原文摘要 · Abstract (English)

Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models. Through trial-and-error, these AI systems typically learn one, near-optimal behavior to solve their tasks. However, there are many use cases in which one would like to assert some level of control, preferably in real time, over how the task is solved. We refer to these modifications of a core task as styles. We combine universal value function approximators (UVFAs) with carefully selected training scenarios, learning algorithms, and data augmentation to create a framework for coaching agents that exhibit styles in complex domains. We demonstrate the framework's application in the AAA video games Horizon Forbidden West and Gran Turismo, and in an open-source humanoid test domain. Despite the different nature of the domains -- car racing, stylized game combat, and humanoid walking -- each agent shows strong coherence to the style requests while still satisfying the main task in its domain. Importantly, the techniques outlined in this paper allow an end user to choose the final behavior at run time, giving them flexible control over the final executed performance.

强化学习游戏AI风格控制可教练性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。