arXiv:2607.23124cs.AIcs.CL2026-07

构建统一框架提升智能体在多场景下的综合能力

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

论文配图:AgentOmnia: Scaling Agentic Models for Full-Scenario Applications
图 1 · 摘自论文原文
  • 设计跨领域任务分类体系与环境合成方法
  • 在多个基准上将通过率从9.16%提升至41.69%
  • 支持自进化,适合工业级智能体研发团队

大语言模型智能体发展迅速,但各领域、能力、任务难度和交互方式之间仍存在割裂。本文提出AgentOmnia框架,实现面向ToC、ToB、ToE三大应用方向的全场景智能体规模化。通过构建域×能力×原子难度的可扩展分类体系,配合OmniaBench进行细粒度诊断。该框架融合双向环境-任务合成、工具依赖、程序化与求解器驱动的流水线,生成5,018个有状态环境、255,375个工具和52,361个任务。利用程序、求解器与验证器提供正确性信号,结合监督微调、在线强化学习与回滚课程完成后训练。评估失败转化为产品需求文档(PRD),驱动自演化。基于Qwen3-30B-A3B-Thinking-2507,OmniaBench挑战集通过率由9.16%提升至37.11%,四基准宏平均通过率从22.86%升至41.69%。在统一协议下领先所有对比基线,并超越Qwen3-235B-A22B-Thinking-2507和Qwen3.5-35B-A3B。性能提升覆盖三类应用场景、十类能力维度、八类原子难度因子及90个一级领域的76个,体现广泛适用性。单轮研究初步验证了PRD引导自演化的可行性,为大规模工业应用铺路。

原文摘要 · Abstract (English)

Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-scenario agentic scaling and present AgentOmnia, a framework coordinating task-space definition, data synthesis, post-training, evaluation, and improvement across To-Consumer (ToC), To-Business (ToB), and To-Employee (ToE) applications. An extensible Domain x Capability x Atomic Difficulty taxonomy aligns these stages and enables fine-grained diagnosis with OmniaBench. AgentOmnia combines bidirectional environment-task synthesis with tool-dependency, program-structured, and solver-based pipelines, constructing 5,018 stateful environments with 255,375 tools and 52,361 tasks. Programs, solvers, and verifiers provide correctness signals, while supervised fine-tuning, online agentic reinforcement learning, and a rollback curriculum support post-training. Evaluation failures translate into Product Requirement Documents (PRDs) for targeted self-evolution. Starting from Qwen3-30B-A3B-Thinking-2507, AgentOmnia raises the pass rate on the OmniaBench challenging subset from 9.16% to 37.11% and the macro-average across OmniaBench, $τ^2$-Bench, DeepPlanning, and VitaBench from 22.86% to 41.69%. Under a unified protocol,it leads the evaluated agentic post-trained baselines on OmniaBench and retains the highest four-benchmark macro-average. It also surpasses Qwen3-235B-A22B-Thinking-2507 on all four benchmarks and exceeds Qwen3.5-35B-A3B on the macro-average. Gains span three application splits, ten capability dimensions, eight atomic-difficulty factors, and 76 of 90 level-1 domains, indicating broad rather than category-specific improvement. A one-round study provides initial evidence for PRD-guided self-evolution, motivating validation at larger scales and in industrial settings.

智能体框架自进化多场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。