构建可扩展的多维度虚拟代理评测基准,突破传统测试局限
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
- 自动生成图结构任务,支持复杂度可控的子任务组合
- 涵盖20个场景36000个任务,人类接受率达91%
- 提供10项能力评估,适合研究多模态智能体性能
随着多模态大语言模型(MLLM)的发展,基于MLLM的虚拟代理展现出卓越性能。然而,现有评测基准存在任务复杂度不可控、人工标注耗时且场景有限、缺乏多维度评估等问题。为此,我们提出OmniBench,一个自动生成、跨平台、基于图结构的评测基准,通过子任务组合实现可控复杂度的任务合成。为评估虚拟代理在图结构上的多元能力,我们进一步提出OmniEval,包含子任务级评估、图级指标以及覆盖10项能力的综合测试。合成数据集包含36,000个图结构任务,覆盖20个场景,人类接受率达91%。在该数据上训练显示,其引导代理的效率优于人工标注数据。我们对多个开源与闭源模型进行多维度评估,揭示其在各类能力上的表现差异,为未来研究铺平道路。项目地址:https://omni-bench.github.io/
原文摘要 · Abstract (English)
As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations, including uncontrollable task complexity, extensive manual annotation with limited scenarios, and a lack of multidimensional evaluation. In response to these challenges, we introduce OmniBench, a self-generating, cross-platform, graph-based benchmark with an automated pipeline for synthesizing tasks of controllable complexity through subtask composition. To evaluate the diverse capabilities of virtual agents on the graph, we further present OmniEval, a multidimensional evaluation framework that includes subtask-level evaluation, graph-based metrics, and comprehensive tests across 10 capabilities. Our synthesized dataset contains 36k graph-structured tasks across 20 scenarios, achieving a 91\% human acceptance rate. Training on our graph-structured data shows that it can more efficiently guide agents compared to manually annotated data. We conduct multidimensional evaluations for various open-source and closed-source models, revealing their performance across various capabilities and paving the way for future advancements. Our project is available at https://omni-bench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。