用在线强化学习动态调整评估函数权重,提升实时战略游戏响应能力。
Online Reinforcement Learning-Based Dynamic Adaptive Evaluation Function for Real-Time Strategy Tasks
- 基于在线强化学习,动态调节评估函数权重以适应战场变化。
- 在多种算法和地图上显著提升评分,大地图下优势更明显。
- 计算开销增加低于6%,适合高实时性策略任务应用。
实时战略任务的有效评估需具备应对动态不可预测环境的自适应机制。本文提出一种基于在线强化学习的动态权重调整方法,用于提升实时战略游戏中对战场局势变化的实时响应能力。在传统静态评估函数基础上,采用梯度下降法实现权重的在线动态更新,并引入权重衰减技术保障稳定性。同时集成AdamW优化器,实时调节强化学习的学习率与衰减率,降低对人工参数调优的依赖。轮对竞赛实验表明,该方法显著提升了兰彻斯特作战模型评估函数、简单评估函数及简单平方根评估函数在IDABCD、IDRTMinimax和组合式AI等规划算法中的应用效果。随着地图规模增大,性能提升更加显著。此外,该方法带来的评估函数计算时间增加均低于6%。所提出的动态自适应评估函数为实时战略任务评估提供了有前景的解决方案。
原文摘要 · Abstract (English)
Effective evaluation of real-time strategy tasks requires adaptive mechanisms to cope with dynamic and unpredictable environments. This study proposes a method to improve evaluation functions for real-time responsiveness to battle-field situation changes, utilizing an online reinforcement learning-based dynam-ic weight adjustment mechanism within the real-time strategy game. Building on traditional static evaluation functions, the method employs gradient descent in online reinforcement learning to update weights dynamically, incorporating weight decay techniques to ensure stability. Additionally, the AdamW optimizer is integrated to adjust the learning rate and decay rate of online reinforcement learning in real time, further reducing the dependency on manual parameter tun-ing. Round-robin competition experiments demonstrate that this method signifi-cantly enhances the application effectiveness of the Lanchester combat model evaluation function, Simple evaluation function, and Simple Sqrt evaluation function in planning algorithms including IDABCD, IDRTMinimax, and Port-folio AI. The method achieves a notable improvement in scores, with the en-hancement becoming more pronounced as the map size increases. Furthermore, the increase in evaluation function computation time induced by this method is kept below 6% for all evaluation functions and planning algorithms. The pro-posed dynamic adaptive evaluation function demonstrates a promising approach for real-time strategy task evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。