arXiv:2607.00190cs.LGcs.AI2026-07

用潜在空间生成反事实走法,帮玩家从输到赢逐步改进策略。

Play Like Champions: Counterfactual Feedback Generation in Latent Space

论文配图:Play Like Champions: Counterfactual Feedback Generation in Latent Space
图 1 · 摘自论文原文
  • 在学习的潜在空间中构建反事实路径,模拟从失败到胜利的改进过程。
  • 基于23,305场职业比赛数据训练模型,在业余数据上验证四类路径生成方法。
  • 提供多粒度反馈,适配不同水平玩家的提升需求,推动人机协同训练。

强化学习已使智能体在众多竞技游戏中超越人类水平。然而,现有研究多聚焦于击败人类而非提供有效反馈,导致难以帮助人类玩家提升。与国际象棋和围棋不同,实时战略游戏(如《星际争霸II》)尚无系统化框架将专家知识转化为可操作建议。本文提出「表现潜空间」(Latent Maps of Performance),利用23,305场职业赛事回放数据训练引导变分自编码器,实现从失败到胜利的潜在空间反事实遍历。设计并验证了四种路径生成策略:线性插值、迭代最优传输、密度正则化梯度上升与神经流匹配,均能在保持专家行为一致性的前提下,生成多步改进轨迹。通过在随机采样的业余玩家回放数据上测试,提取多层次反馈,支持不同阶段玩家的成长。结果表明路径搜索方法存在权衡,未来应更关注模型驱动的人类提升解决方案。

原文摘要 · Abstract (English)

Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games. As a byproduct, researchers have begun studying how these agents play, extracting behavioral representations, analyzing decision structure, and modeling the latent geometry of expert performance. However, this growing body of work has overwhelmingly focused on defeating human players rather than providing feedback, leaving a critical gap in creating model solutions to improve human players. Unlike chess and Go, where AI has become integral to player training, real-time strategy (RTS) games lack principled frameworks for translating expert knowledge into actionable feedback. We introduce Latent Maps of Performance, a framework for counterfactual path generation. We focus on StarCraft~II data to model player improvement as an algorithmic recourse within a learned representation space. As inspiration for our work, we have looked at the championship model used in sports science. We trained a Guided Variational Autoencoder model on 23,305 professional tournament replays, enabling counterfactual traversal between losing and winning gameplay profiles. To fulfill our goal, we have devised and verified four traversal strategies on out-of-distribution (OOD) data randomly sampled from a dataset of amateur replays, namely linear interpolation, iterative optimal transport, density-regularized gradient ascent, and neural flow matching, each designed to generate multi-step improvement trajectories that remain grounded in observed expert behavior while moving a player's profile toward winning configurations. Feedback is extracted at multiple granularities to support players at different stages of improvement. Finally, we conclude that there is a trade-off between the path-finding methods we employ and hope that future research will focus on developing model solutions for human improvement.

强化学习游戏AI反事实推理人类增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。