arXiv:2512.08139cs.LG2025-12被引 1

训练能应对未知环境的智能体,提升AI在复杂世界中的适应力。

Robust Agents in Open-Ended Worlds

  • 用程序生成环境和对抗性课程,强化智能体泛化能力
  • 在双人零和游戏中显著提升智能体鲁棒性,跨场景表现更优
  • 适用于需要高适应性的AI系统,如游戏、机器人或大模型防御

人工智能在各类应用中日益普及,亟需能够适应不断变化的开放世界并保持稳健表现的智能体。本论文聚焦于训练与评估具备强泛化能力的智能体,使其在未见过的环境、分布外输入以及与其他智能体交互时仍能有效运作。研究首先提出MiniHack——基于NetHack的游戏框架,通过程序化内容生成构建多样化强化学习任务,用于测试泛化性能。接着提出Maestro方法,利用对抗性课程逐步提升双人零和博弈中智能体的鲁棒性与通用性。进一步采用质量-多样性方法,在足球类复杂视频游戏(具有合作与竞争交织特性)中系统识别先进预训练强化学习策略的脆弱点。最后将研究扩展至大语言模型领域,通过进化搜索生成多样化的恶意提示,诊断并增强大模型对对抗性输入的鲁棒性。该工作为未来智能体鲁棒性发展奠定基础,推动其在不可预测挑战中持续稳定运行。

原文摘要 · Abstract (English)

The growing prevalence of artificial intelligence (AI) in various applications underscores the need for agents that can successfully navigate and adapt to an ever-changing, open-ended world. A key challenge is ensuring these AI agents are robust, excelling not only in familiar settings observed during training but also effectively generalising to previously unseen and varied scenarios. In this thesis, we harness methodologies from open-endedness and multi-agent learning to train and evaluate robust AI agents capable of generalising to novel environments, out-of-distribution inputs, and interactions with other co-player agents. We begin by introducing MiniHack, a sandbox framework for creating diverse environments through procedural content generation. Based on the game of NetHack, MiniHack enables the construction of new tasks for reinforcement learning (RL) agents with a focus on generalisation. We then present Maestro, a novel approach for generating adversarial curricula that progressively enhance the robustness and generality of RL agents in two-player zero-sum games. We further probe robustness in multi-agent domains, utilising quality-diversity methods to systematically identify vulnerabilities in state-of-the-art, pre-trained RL policies within the complex video game football domain, characterised by intertwined cooperative and competitive dynamics. Finally, we extend our exploration of robustness to the domain of LLMs. Here, our focus is on diagnosing and enhancing the robustness of LLMs against adversarial prompts, employing evolutionary search to generate a diverse range of effective inputs that aim to elicit undesirable outputs from an LLM. This work collectively paves the way for future advancements in AI robustness, enabling the development of agents that not only adapt to an ever-evolving world but also thrive in the face of unforeseen challenges and interactions.

智能体鲁棒性强化学习大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。