arXiv:2506.16995cs.AI2025-06中稿 · Frontiers of Compu…

让游戏智能体既强又多样,保留个性打法的同时提升实力。

Policy Improvement with Style-Specific Demonstrations

  • 用混合策略优化统一在线与离线样本损失,保持风格差异。
  • 在多个环境中表现优于纯在线算法,同时保留原始玩法风格。
  • 适合想提升游戏智能体多样性与性能的开发者和研究者。

具备多样玩法风格的高超游戏智能体能丰富游戏体验并提升重玩价值。然而,当前基于强化学习的游戏AI主要关注提升能力,而基于进化算法的方法虽能生成多样风格,但性能远低于强化学习方法。为此,本文提出混合近端策略优化(MPPO),旨在提升现有低效智能体的能力,同时保留其独特玩法风格。MPPO统一了在线与离线样本的损失目标,并引入隐式约束,通过调整样本的经验分布来逼近示范策略。在不同规模环境中的实验表明,MPPO在能力上达到甚至超过纯在线算法水平,同时有效保留了示范者的玩法风格。该方法为生成高效且多样的游戏智能体提供了有效路径,有助于实现更富吸引力的游戏体验。

原文摘要 · Abstract (English)

Proficient game agents with diverse play styles enrich the gaming experience and enhance the replay value of games. However, recent advancements in game AI based on reinforcement learning have predominantly focused on improving proficiency, whereas methods based on evolution algorithms generate agents with diverse play styles but exhibit subpar performance compared to RL methods. To address this gap, this paper proposes Mixed Proximal Policy Optimization (MPPO), a method designed to improve the proficiency of existing suboptimal agents while retaining their distinct styles. MPPO unifies loss objectives for both online and offline samples and introduces an implicit constraint to approximate demonstrator policies by adjusting the empirical distribution of samples. Empirical results across environments of varying scales demonstrate that MPPO achieves proficiency levels comparable to, or even superior to, pure online algorithms while preserving demonstrators' play styles. This work presents an effective approach for generating highly proficient and diverse game agents, ultimately contributing to more engaging gameplay experiences.

游戏AI策略优化多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。