让智能体适应未知环境,关键在扩大规则多样性而非只堆数据。
Scalable Environments Drive Generalizable Agents

- 用可扩展的规则集构建环境,突破固定任务限制。
- 新分类法区分轨迹、任务与环境三类扩展方式。
- 适合研究通用智能体与鲁棒性强化学习的团队。
通用智能体需适应训练分布外的多样任务与未知环境。本文主张,真正泛化依赖于环境扩展:即扩大智能体交互的可执行规则集分布,而非仅增加固定基准下的轨迹或任务数量。当前方法多聚焦于收集更多经验或扩展任务集,却忽视了界面、动态、观测或反馈信号变化带来的脆弱性。核心挑战是世界级分布偏移:智能体必须系统接触规则本质不同的环境。为此,我们提出统一分类法,按主要产出和规则集变化区分轨迹扩展、任务扩展与环境扩展。基于此,我们综述了可扩展环境的构建范式,对比了强调可控性与可验证性的程序生成器,以及提供更广覆盖与开放性的生成世界模型。进一步探讨了环境扩展与状态化学习机制的结合,强调跨环境自适应的可学习更新规则。最后讨论替代视角,认为可扩展环境是实现稳健通用智能体可测量、可控制进步的基础。
原文摘要 · Abstract (English)
Generalizable agents should adapt to diverse tasks and unseen environments beyond their training distribution. This position paper argues that such generalization requires environment scaling: expanding the distribution of executable rule-sets that agents interact with, rather than only increasing trajectories or tasks within fixed benchmarks. Current scaling practices largely focus on collecting more experience or broader task sets under fixed interaction rules, leaving agents brittle when underlying interfaces, dynamics, observations, or feedback signals change. The core challenge is therefore a world-level distribution shift: agents need systematic exposure to environments with meaningfully different executable rule-sets. To clarify this challenge, we propose a unified taxonomy that separates trajectory scaling, task scaling, and environment scaling by their primary deliverables and by what changes in the executable rule-set. Building on this taxonomy, we synthesize construction paradigms for scalable environments, contrasting programmatic generators that prioritize controllability and verifiability with generative world models that offer broader coverage and open-endedness. We further outline how environment scaling can be coupled with stateful learning mechanisms, emphasizing learned update rules for cross-environment adaptation. We conclude by discussing alternative perspectives and argue that scalable environments provide the essential substrate for measurable and controllable progress toward robust general agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。