arXiv:2606.06560cs.LGcs.AI2026-06中稿 · the Second Worksho…

MacArena为macOS设计新基准,揭示现有模型在苹果生态下表现大幅下滑。

MacArena: Benchmarking Computer Use Agents on an Online macOS Environment

论文配图:MacArena: Benchmarking Computer Use Agents on an Online macOS Environment
图 1 · 摘自论文原文
  • 构建421个经人工验证的macOS任务,运行于Apple Silicon原生虚拟环境。
  • 主流模型在MacArena上表现比原有基准下降超26%,凸显系统差异挑战。
  • 适合评估跨平台GUI智能体真实能力,尤其关注苹果生态开发者。

计算机使用智能体(CUAs)通过视觉与控制原语操作图形用户界面(GUI),其能力迅速提升,部分得益于如OSWorld等标准化在线评估基准,这些基准既用于评估也用于强化学习训练。然而,macOS在该领域仍被忽视:现有唯一基准macOSWorld仅覆盖少数第一方应用,任务较简单,且运行于x86虚拟机,不兼容Apple Silicon。我们提出MacArena,一个包含421个手动验证任务、涵盖50个应用的基准,融合了经筛选的OSWorld任务、来自macOSWorld的内容以及49个全新macOS原生任务,全部在Apple Silicon的原生Virtualization框架上运行。我们认为macOS带来超越基于Linux基准的独特GUI挑战,我们的评估支持这一观点:当前模型在已有基准上的优异表现可能反映对任务分布的熟悉度,而非真正的跨平台GUI能力。值得注意的是,模型在移植任务与macOS原生任务上的排名完全反转,领先模型在MacArena子集上落后超过26%,表明当前GUI智能体在macOS环境中面临真正更严峻的挑战。

原文摘要 · Abstract (English)

Computer-use agents (CUAs) operate graphical user interfaces (GUIs) through vision and control primitives, and their capabilities have advanced rapidly, driven in part by standardized online evaluation benchmarks such as OSWorld, which serve both as evaluation tools and as training environments for reinforcement learning. However, macOS remains underserved in this landscape: the only existing benchmark, macOSWorld, covers a narrow slice of first-party applications with simpler tasks, and runs on x86 virtual machines incompatible with Apple Silicon. We introduce MacArena, a benchmark of 421 manually verified tasks spanning 50 applications that combines a curated port of OSWorld tasks, content sourced from macOSWorld, and 49 new macOS-native tasks, all running on Apple's native Virtualization framework on Apple Silicon. We argue that macOS presents distinct GUI challenges beyond what Linux-based benchmarks capture, and our evaluation supports this claim: strong model performance on existing benchmarks can reflect familiarity with task distributions rather than genuine cross-platform GUI competence. Notably, model rankings invert between ported and macOS-native tasks, with a leading model trailing by over 26% on the MacArena subset, suggesting that macOS poses a genuinely harder environment for current GUI agents.

GUI智能体macOS基准Apple Silicon评估挑战

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。