构建可扩展的异构多智能体对抗强化学习框架,支持高保真仿真训练。
A Framework for Scalable Heterogeneous Multi-Agent Adversarial Reinforcement Learning in IsaacLab
- 基于HAPPO算法设计异构对抗MARL训练机制。
- 在多个基准场景中实现高效策略训练与高吞吐仿真。
- 适合机器人对抗任务研究者快速实验与验证。
多智能体强化学习(MARL)是动态环境中机器人协作的核心。尽管以往研究集中于合作场景,但对抗性交互在追逃、安防和竞争操控等真实应用中同样关键。本文扩展IsaacLab框架,支持在高保真物理仿真中可扩展地训练对抗策略。引入一套包含异构智能体、目标与能力不对称的对抗式MARL环境。平台集成改进版异构智能体强化学习(HAPPO),结合近端策略优化,实现对抗动态下的高效训练与评估。在多个基准场景中的实验表明,该框架能有效建模并训练形态多样、具备鲁棒性的多智能体竞争策略,同时保持高吞吐与仿真真实性。代码与基准数据集开源:https://github.com/DIRECTLab/IsaacLab-HARL。
原文摘要 · Abstract (English)
Multi-Agent Reinforcement Learning (MARL) is central to robotic systems cooperating in dynamic environments. While prior work has focused on these collaborative settings, adversarial interactions are equally critical for real-world applications such as pursuit-evasion, security, and competitive manipulation. In this work, we extend the IsaacLab framework to support scalable training of adversarial policies in high-fidelity physics simulations. We introduce a suite of adversarial MARL environments featuring heterogeneous agents with asymmetric goals and capabilities. Our platform integrates a competitive variant of Heterogeneous Agent Reinforcement Learning with Proximal Policy Optimization (HAPPO), enabling efficient training and evaluation under adversarial dynamics. Experiments across several benchmark scenarios demonstrate the framework's ability to model and train robust policies for morphologically diverse multi-agent competition while maintaining high throughput and simulation realism. Code and benchmarks are available at: https://github.com/DIRECTLab/IsaacLab-HARL .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。