构建新基准评估电脑操作智能体的广度与深度能力。
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
- 按自动化等级和用户需求层级设计双维评估体系
- 顶尖视觉语言模型代理仍难完成高阶感知推理任务
- 适合研究者提升智能体实用性与落地可行性
计算机操作智能体在提升人类生产力和开拓新应用形态方面展现出巨大潜力。然而,现有基准未能充分反映任务内部异质性、智能体能力差异及其与真实用户需求的对齐,制约了针对性能力发展和研究成果向实际部署的转化。为此,我们提出OS-MAP,一个面向日常电脑自动化任务的基准,包含416个真实任务,覆盖15个应用。其核心在于两个维度:五级自动化分类与基于真实用户需求层级的泛化范围。该设计可精细分析所需能力与现实场景的匹配度,形成性能-泛化评估矩阵,实现结构化、全面的评估。实验表明,即使采用最先进的视觉语言模型作为基座的智能体,在涉及感知、推理与协调的高级任务上依然表现不佳,凸显了深入理解当前优势与局限的重要性,以推动未来研究与部署进展。所有代码、环境、基线及数据已公开于https://github.com/OS-Copilot/OS-Map。
原文摘要 · Abstract (English)
Computer-using agents have shown strong potential to boost human productivity and enable new application forms across platforms. While recent advances have led to usable applications, existing benchmarks fail to account for the internal task heterogeneity and the corresponding agent capabilities, as well as their alignment with actual user demands-hindering both targeted capability development and the reliable transition of research progress into practical deployment. To bridge the gap, we present OS-MAP, a benchmark for daily computer-using automation that organizes its 416 realistic tasks across 15 applications along two key dimensions: a five-level taxonomy of automation and a generalization scope derived from a real-world user demand hierarchy. To enable fine-grained analysis of required capabilities and alignment with real-world scenarios, OS-MAP evaluates agents along two dimensions: automation level across a five-level taxonomy, and generalization scope across a demand hierarchy. This design captures varying levels of required agent autonomy and generalization, forming a performance-generalization evaluation matrix for structured and comprehensive assessment. Experiments show that even State-of-the-Art agents with VLM backbones struggle with higher-level tasks involving perception, reasoning, and coordination-highlighting the need for a deeper understanding of current strengths and limitations to drive the future progress in computer-using agents research and deployment. All code, environments, baselines, and data are publicly available at https://github.com/OS-Copilot/OS-Map.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。