用路径点抽象地形,让军事模拟强化学习更高效
Abstracting Geo-specific Terrains to Scale Up Reinforcement Learning
- 用路径点自动生成多层地形抽象表示
- 在新任务中实现更快学习且轨迹接近真人玩家
- 适合需真实地形与复杂目标的军事训练场景
多智能体强化学习(MARL)正广泛应用于交互式仿真中动态自适应虚拟角色的训练,尤其在具有地理特异性的地形上。如Unity的ML-Agents框架使这类实验对仿真社区更易访问。军事训练模拟虽受益于MARL进展,但其复杂的连续性、随机性、部分可观测性、非平稳性及基于条令的特性带来巨大计算需求,且要求使用地理特异性地形,进一步加剧资源压力。本研究利用Unity的路径点(waypoints),自动构建多层级的地形抽象表示,在保持策略跨表示迁移能力的同时,显著提升强化学习可扩展性。初步探索在一种新型MARL场景中进行,双方目标不同,结果表明基于路径点的导航能实现更快、更高效的训练,并生成与《CSGO》真人玩家相似的行动轨迹。该研究揭示了路径点导航在降低军事训练模拟中MARL模型开发与训练成本方面的潜力,尤其适用于需要地理特异性地形和差异化目标的场景。
原文摘要 · Abstract (English)
Multi-agent reinforcement learning (MARL) is increasingly ubiquitous in training dynamic and adaptive synthetic characters for interactive simulations on geo-specific terrains. Frameworks such as Unity's ML-Agents help to make such reinforcement learning experiments more accessible to the simulation community. Military training simulations also benefit from advances in MARL, but they have immense computational requirements due to their complex, continuous, stochastic, partially observable, non-stationary, and doctrine-based nature. Furthermore, these simulations require geo-specific terrains, further exacerbating the computational resources problem. In our research, we leverage Unity's waypoints to automatically generate multi-layered representation abstractions of the geo-specific terrains to scale up reinforcement learning while still allowing the transfer of learned policies between different representations. Our early exploratory results on a novel MARL scenario, where each side has differing objectives, indicate that waypoint-based navigation enables faster and more efficient learning while producing trajectories similar to those taken by expert human players in CSGO gaming environments. This research points out the potential of waypoint-based navigation for reducing the computational costs of developing and training MARL models for military training simulations, where geo-specific terrains and differing objectives are crucial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。