用对比学习提升游戏智能体的状态表征能力
Learning Representations in Video Game Agents with Supervised Contrastive Imitation Learning
- 将监督对比学习引入模仿学习,优化观察到动作的映射关系
- 在3D和2D游戏上实现更快收敛与更好泛化性能
- 适用于任意连续动作空间的游戏智能体训练
本文提出一种将监督对比学习(Supervised Contrastive Learning, SupCon)应用于模仿学习(Imitation Learning, IL)的新方法,旨在提升视频游戏环境中智能体的状态表征质量。目标是学习出能捕捉动作相关因素的潜在表示,从而更准确地建模观察与示范者(如玩家)行为之间的因果关系——例如,当障碍物出现时玩家会跳跃。所提方法将SupCon损失拓展至连续输出空间,使该方法不受环境动作类型限制。在3D游戏Astro Bot、Returnal及多个2D Atari游戏中进行实验,结果表明,相比仅使用监督动作预测损失的基线模型,该方法在状态表示质量、学习收敛速度和泛化能力方面均有显著提升。
原文摘要 · Abstract (English)
This paper introduces a novel application of Supervised Contrastive Learning (SupCon) to Imitation Learning (IL), with a focus on learning more effective state representations for agents in video game environments. The goal is to obtain latent representations of the observations that capture better the action-relevant factors, thereby modeling better the cause-effect relationship from the observations that are mapped to the actions performed by the demonstrator, for example, the player jumps whenever an obstacle appears ahead. We propose an approach to integrate the SupCon loss with continuous output spaces, enabling SupCon to operate without constraints regarding the type of actions of the environment. Experiments on the 3D games Astro Bot and Returnal, and multiple 2D Atari games show improved representation quality, faster learning convergence, and better generalization compared to baseline models trained only with supervised action prediction loss functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。