arXiv:2604.18292cs.AIcs.CL2026-04被引 21

构建可自进化的真实环境,让智能体持续学习并突破能力边界。

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

论文配图:Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
图 1 · 摘自论文原文
  • 自动生成多样化、可验证的任务,支持难度可控的环境合成。
  • 在23个基准上超越主流商用模型,14B版本表现最优。
  • 适合研究通用智能体长期学习与环境协同演化的方向。

大型语言模型正被期待作为与外部状态化工具环境交互的通用智能体。尽管模型上下文协议(MCP)和更广泛的智能体技能提供了连接智能体与可扩展真实服务的统一接口,但训练鲁棒智能体仍受限于缺乏真实环境及长期学习的系统性机制。本文提出 extbf{Agent-World},一个通过可扩展环境推动通用智能体智能进化的自演化训练场。该系统包含两大组件:(1) 智能体环境-任务发现,从数千个真实世界主题数据库和可执行工具生态中自主探索,合成可验证且难度可控的任务;(2) 连续自演化智能体训练,结合多环境强化学习与自演化智能体竞技场,通过动态任务生成识别能力短板并驱动针对性学习,实现智能体策略与环境的共同演化。在23个挑战性智能体基准测试中,Agent-World-8B与14B均持续优于强大多产模型和环境扩展基线。进一步分析揭示了环境多样性与自演化轮次间的缩放规律,为构建通用智能体智能提供新洞见。

原文摘要 · Abstract (English)

Large language models are increasingly expected to serve as general-purpose agents that interact with external, stateful tool environments. The Model Context Protocol (MCP) and broader agent skills offer a unified interface for connecting agents with scalable real-world services, but training robust agents remains limited by the lack of realistic environments and principled mechanisms for life-long learning. In this paper, we present \textbf{Agent-World}, a self-evolving training arena for advancing general agent intelligence through scalable environments. Agent-World has two main components: (1) Agentic Environment-Task Discovery, which autonomously explores topic-aligned databases and executable tool ecosystems from thousands of real-world environment themes and synthesizes verifiable tasks with controllable difficulty; and (2) Continuous Self-Evolving Agent Training, which combines multi-environment reinforcement learning with a self-evolving agent arena that automatically identifies capability gaps through dynamic task synthesis and drives targeted learning, enabling the co-evolution of agent policies and environments. Across 23 challenging agent benchmarks, Agent-World-8B and 14B consistently outperforms strong proprietary models and environment scaling baselines. Further analyses reveal scaling trends in relation to environment diversity and self-evolution rounds, offering insights for building general agent intelligence.

智能体自进化环境合成强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。