评测大模型在办公室任务中以成本为基准的执行能力。
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

- 构建100个真实办公任务,每项平均需2.32小时人工完成。
- 用人工工时和任务价格代理双指标评估成本与价值。
- 适合关注大模型经济效率与实际应用落地的研究者。
大型语言模型(LLM)代理正被期待协助用户完成任务,但现有评测难以判断其在合理成本下完成办公流程的能力。我们提出OmegaUse-OfficeVal,一个针对长周期办公任务、带有任务级经济基准的评测基准。该基准包含100个由从业者提出的办公请求,经隐私保护处理后生成,平均需2.32小时人工完成。每个任务配备两个经济信号:人工工时与任务价格代理,可直接比较人类成本与模型推理成本,并实现价值加权评估。为保障评测稳定性,我们基于细粒度评分标准开发了代码验证器。我们评估了多个前沿大模型及人类基线。尽管所有模型均显著低于人工成本且速度更快,但尚未达到人类水平的交付质量。代码与数据集已完全开源,更多信息见项目官网:https://omegause-officeval.github.io。
原文摘要 · Abstract (English)
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, and more information is available on our project website: https://omegause-officeval.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。