构建跨场景通用智能体评估基准,揭示模型在复杂任务中的真实短板。
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

- 基于90个一级领域构建分层场景体系,覆盖ToC/ToB/ToE多类应用
- 包含1431个可执行任务,挑战当前前沿模型通过率不足60%
- 细粒度分析能力维度与困难因子,适合评估智能体系统性能力
大型语言模型正从文本生成器演变为能理解用户需求、调用外部工具并通过交互完成复杂任务的通用智能体。然而现有评测基准多局限于特定场景、工具生态或交互形式,难以系统刻画模型在异构应用环境中的能力。本文提出OmniaBench,一个面向多样化场景的通用智能体评测基准,具有显式状态空间设计。通过应用商店、产品文档、行业资源、网络检索与人工精修,构建涵盖ToC、ToB、ToE的层级化分类体系,共90个一级与354个二级领域。基于此,通过四种互补路径(DAG、DAG-S、Solver、Program)构建可执行环境,合成单轮与多轮任务。该基准引入十维能力分类与八种组合式基础难度因子,支持精细化评估。数据集含1,431个任务,另设644个高挑战子集以降低评估成本并缓解公开后污染风险。测试显示,即便前沿模型Claude-Sonnet-5和GPT-5.6-Sol的Overall Pass@1得分也仅分别为58.54%与57.14%。进一步分析揭示各领域与能力间显著差异,以及规划、约束保持与自适应修正等持续性短板。OmniaBench为刻画通用智能体的能力边界提供了广度与诊断性兼具的评测体系。
原文摘要 · Abstract (English)
Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tasks through interaction. However, existing agent benchmarks often focus on limited scenarios, tool ecosystems, or interaction formats, making it difficult to systematically characterize model capabilities across heterogeneous application settings. We introduce OmniaBench, a benchmark for evaluating general agents across diverse scenarios with explicit state spaces. We derive application-oriented scenario knowledge from app stores, product documents, industry resources, Web retrieval, and human refinement, forming a hierarchical taxonomy that spans ToC, ToB and ToE with 90 level-1 and 354 level-2 domains. Based on this taxonomy, we construct executable environments and synthesize single-turn and multi-turn tasks through four complementary routes: DAG, DAG-S, Solver, and Program. OmniaBench further introduces a ten-dimensional capability taxonomy and eight compositional atomic difficulty factors to support fine-grained evaluation and analysis. The resulting dataset contains 1,431 tasks, together with a challenging subset of 644 tasks designed to reduce evaluation cost and mitigate potential contamination of the full set after public release. The bench presents substantial challenges to current frontier models, with even Claude-Sonnet-5 and GPT-5.6-Sol achieving Overall Pass@1 scores of only 58.54 and 57.14, respectively. Further analyses reveal clear differences across domains and capabilities, as well as persistent limitations in planning, constraint maintenance, and adaptive correction. OmniaBench provides a broad and diagnostic benchmark for characterizing the capability boundaries of general agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。