用对抗演化自动生成复杂环境,让智能体在不断进化的挑战中提升策略能力。
Co-Evolving Complexity: An Adversarial Framework for Automatic MARL Curricula
- 让攻击者与防御者互为对手,动态生成针对防御弱点的新关卡。
- 仅需少量训练即涌现出包抄、掩护等复杂战术行为。
- 适合研究自适应训练环境、多智能体强化学习的学者参考。
通用智能体的发展与其训练环境密不可分。尽管模型和数据规模的扩展已带来显著能力提升,但环境的复杂性、多样性和交互性仍是一大瓶颈。人工设计的环境有限且常含隐式偏差,限制了智能体发展真正泛化和鲁棒的能力。本文提出一种通过对抗博弈生成无限自适应课程的范式:一组协作的多智能体防御者需应对一个程序化生成的攻击者。攻击者学习生成更具挑战性的敌方单位配置,动态创造针对性的新世界;防御者团队则同步学习协作策略以克服这些威胁。这种共演化机制使环境复杂性自发增长,提供持续不断的新颖且相关的训练数据。实验表明,仅经少量训练,系统便涌现出攻击者的包抄与掩护、防御者的集中攻击与分散部署等复杂智能行为。结果表明,对抗共演化是自动扩展环境复杂性的有力机制,可有效推动智能体向更高鲁棒性与战略深度演进。
原文摘要 · Abstract (English)
The advancement of general-purpose intelligent agents is intrinsically linked to the environments in which they are trained. While scaling models and datasets has yielded remarkable capabilities, scaling the complexity, diversity, and interactivity of environments remains a crucial bottleneck. Hand-crafted environments are finite and often contain implicit biases, limiting the potential for agents to develop truly generalizable and robust skills. In this work, we propose a paradigm for generating a boundless and adaptive curriculum of challenges by framing the environment generation process as an adversarial game. We introduce a system where a team of cooperative multi-agent defenders learns to survive against a procedurally generative attacker. The attacker agent learns to produce increasingly challenging configurations of enemy units, dynamically creating novel worlds tailored to exploit the defenders' current weaknesses. Concurrently, the defender team learns cooperative strategies to overcome these generated threats. This co-evolutionary dynamic creates a self-scaling environment where complexity arises organically from the adversarial interaction, providing an effectively infinite stream of novel and relevant training data. We demonstrate that with minimal training, this approach leads to the emergence of complex, intelligent behaviors, such as flanking and shielding by the attacker, and focus-fire and spreading by the defenders. Our findings suggest that adversarial co-evolution is a powerful mechanism for automatically scaling environmental complexity, driving agents towards greater robustness and strategic depth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。