arXiv:2509.21234cs.LGcs.MA2025-09被引 1

让强化学习环境动态变化,测试智能体实时适应能力

AbideGym: Turning Static RL Worlds into Adaptive Challenges

  • 用动态扰动机制让训练环境实时变化
  • 揭示静态策略在变化中失败的弱点
  • 适合研究持续学习与鲁棒泛化的研究者

基于强化学习训练的智能体常因环境动态变化而失效,这一问题在静态基准测试中尤为突出。AbideGym 是一个动态的 MiniGrid 封装环境,引入了智能体感知的扰动机制和可扩展的复杂度设计,旨在促进智能体在单个实验过程中实现自适应。该框架通过暴露静态策略的脆弱性,推动模型在持续学习、课程学习和鲁棒泛化方面的进步,提供了一个模块化且可复现的评估体系。

原文摘要 · Abstract (English)

Agents trained with reinforcement learning often develop brittle policies that fail when dynamics shift, a problem amplified by static benchmarks. AbideGym, a dynamic MiniGrid wrapper, introduces agent-aware perturbations and scalable complexity to enforce intra-episode adaptation. By exposing weaknesses in static policies and promoting resilience, AbideGym provides a modular, reproducible evaluation framework for advancing research in curriculum learning, continual learning, and robust generalization.

强化学习动态环境持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。