arXiv:2512.04987cs.CL2025-12被引 16

构建复杂交互环境,训练出更强的自主智能体模型

Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction

  • 用三维度框架自动构建多样复杂的交互环境
  • 在SWE-bench等任务上超越主流开源模型
  • 适合研究自主智能体与大规模环境构建的学者

大型语言模型从被动响应者向自主智能体的演进,要求学习范式从静态模仿转向激励驱动的决策。然而,这一转变受限于缺乏可扩展的基础设施来构建高质量的交互信号以支持策略学习。为此,我们提出一套系统化方法,全面扩展交互环境的多样性与复杂性。该方法从三个独立维度入手:(1) 复杂性:NexAU 是一个灵活的智能体框架,通过简单配置即可构建复杂智能体层级;(2) 多样性:NexA4A 能从自然语言自动生成多样化的智能体层级,覆盖无限领域;(3) 真实性:NexGAP 通过整合动态真实世界环境,弥合仿真与现实之间的差距,实现有根基的轨迹生成。我们在该基础设施构建的多样化、复杂化交互环境中训练 Nex-N1 模型。在 SWE-bench 与 tau2 等基准测试中,Nex-N1 持续优于当前最先进开源模型,并在复杂智能体任务上达到与前沿闭源模型相当的性能。我们已开源 Nex 生态系统及模型权重,以推动后续研究。

原文摘要 · Abstract (English)

The evolution of Large Language Models (LLMs) from passive responders to autonomous agents necessitates a fundamental shift in learning paradigms -- from static imitation to incentive-driven decision making. However, this transition is significantly impeded by the lack of scalable infrastructure capable of constructing high-quality interaction signals for effective policy learning. To address this, we introduce a comprehensive method designed to systematically scale the diversity and complexity of interactive environments. Our method realizes this scaling by addressing three orthogonal dimensions: (1) Complexity: NexAU, a flexible agent framework that supports building complex agent hierarchies via simple configurations; (2) Diversity: NexA4A automatically generates diverse agent hierarchies from natural language to cover infinite domains; and (3) Fidelity: NexGAP bridges the simulation-reality gap by integrating dynamic real-world environment for grounded trajectories synthesis. We train Nex-N1 upon the diverse and complex interactive environments established by our infrastructure. Empirical results on benchmarks such as SWE-bench and tau2 demonstrate that Nex-N1 consistently outperforms SOTA open-source models and achieves competitive performance against frontier proprietary models on complex agentic tasks. We open-source the Nex ecosystem and model weights to facilitate further research.

智能体环境构建大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。