评测AI在复杂文件依赖场景下的工作空间学习能力,发现当前模型表现远低于人类。
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies

- 构建包含2万+文件的逼真工作空间,模拟真实跨文件依赖关系。
- 现有最佳模型仅达60%准确率,远低于人类80.7%的水平。
- 提供轻量版基准,可降低70%评估成本,适合研究者快速测试。
工作空间学习要求AI代理识别、推理、利用并更新工作者环境中异构文件间的显式与隐式依赖关系,从而有效完成常规及高级任务。尽管其重要性突出,现有基准大多基于预设或合成文件,缺乏真实世界依赖关系,导致工作空间层面评估严重不足。为此,我们提出Workspace-Bench,一个面向大规模文件依赖的工作空间学习评测基准。该基准构建了5个用户档案、74种文件类型、共20,476个文件(最大达20GB),并设计388项任务,每项任务配有独立的文件依赖图,需在7,399个评价维度上完成跨文件检索、上下文推理与自适应决策。我们还推出Workspace-Bench-Lite,一个100任务子集,在保留分布特征的同时将评估成本降低约70%。我们对4种主流代理框架和7个基础模型进行了评测,结果表明当前代理在工作空间学习上仍不可靠:最佳表现仅约60%,显著低于人类80.7%的水平,平均性能仅为43.3%。
原文摘要 · Abstract (English)
Workspace learning requires AI agents to identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a worker's workspace, enabling them to complete both routine and advanced tasks effectively. Despite its importance, existing relevant benchmarks largely evaluate agents on pre-specified or synthesized files with limited real-world dependencies, leaving workspace-level evaluation underexplored. To this end, we introduce Workspace-Bench, a benchmark for evaluating AI agents on Workspace Learning involving Large-Scale File Dependencies. We construct realistic workspaces with 5 worker profiles, 74 file types, 20,476 files (up to 20GB) and curate 388 tasks, each with its own file dependency graph, evaluated across 7,399 total rubrics that require cross-file retrieval, contextual reasoning, and adaptive decision-making. We further provide Workspace-Bench-Lite, a 100-task subset that preserves the benchmark distribution while reducing evaluation costs by about 70%. We evaluate 4 popular agent harnesses and 7 foundation models. Experimental results show that current agents remain far from reliable workspace learning, where the best reaches only about 60%, substantially below the human result of 80.7%, and the average performance across agents is only 43.3%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。