arXiv:2601.02399cs.SEcs.AI2026-01被引 5

评测多模态智能体在专业软件中的能力,发现当前最强模型仅24.4%成功率。

ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments

论文配图:ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
图 1 · 摘自论文原文
  • 构建专业软件任务层级体系,覆盖6大领域13款核心应用
  • 真实计算机环境执行,最先进代理在复杂任务中成功率仅24.4%
  • 融合人工评估,为专业场景智能体设计提供关键参考

多模态智能体在通用计算机任务上进展迅速,但现有基准仍局限于浏览器和基础桌面应用,难以覆盖主导科研与工业实践的专业软件流程。为此,我们提出ProSoftArena,一个专用于评估多模态智能体在专业软件环境中表现的基准与平台。建立首个面向专业软件使用的能力建构层级,构建涵盖6个学科、13种核心专业应用的436个真实工作与研究任务。为确保评估可靠可复现,搭建可执行的真实计算机环境,采用基于执行的评估框架,并首次引入人机协同评估范式。大量实验表明,即使最优智能体在L2级任务中成功率也仅达24.4%,对L3级跨软件流程任务完全失败。深入分析揭示当前智能体局限,为未来更高效的设计原则提供洞见,推动专业软件场景下智能体能力提升。项目开源地址:https://prosoftarena.github.io。

原文摘要 · Abstract (English)

Multimodal agents are making rapid progress on general computer-use tasks, yet existing benchmarks remain largely confined to browsers and basic desktop applications, falling short in professional software workflows that dominate real-world scientific and industrial practice. To close this gap, we introduce ProSoftArena, a benchmark and platform specifically for evaluating multimodal agents in professional software environments. We establish the first capability hierarchy tailored to agent use of professional software and construct a benchmark of 436 realistic work and research tasks spanning 6 disciplines and 13 core professional applications. To ensure reliable and reproducible assessment, we build an executable real-computer environment with an execution-based evaluation framework and uniquely incorporate a human-in-the-loop evaluation paradigm. Extensive experiments show that even the best-performing agent attains only a 24.4\% success rate on L2 tasks and completely fails on L3 multi-software workflow. In-depth analysis further provides valuable insights for addressing current agent limitations and more effective design principles, paving the way to build more capable agents in professional software settings. This project is available at: https://prosoftarena.github.io.

多模态智能体专业软件基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。