为真实业务场景中的AI代理设计了94个评估任务,更贴近实际应用。
AlphaEval: Evaluating Agents in Production

- 从7家公司的实际需求中提取任务,构建生产级评估基准。
- 涵盖6个职业领域,评估完整代理系统而非单一模型性能。
- 提供可复用的流程框架,帮助企业快速搭建自定义评估体系。
AI代理在商业环境中的快速部署已超越评估方法的发展,现有基准多基于事后整理的任务和确定性指标,与生产环境差异显著:需求常含隐含约束,输入为异构多模态文档且信息分散,任务需未明示的专业知识,输出为长周期专业成果,成功由随时间演化的专家标准判断。我们提出AlphaEval,一个源自7家公司在核心业务中部署AI代理的真实任务集合,共94项,覆盖6个O*NET职业领域。不同于以模型为中心的基准,AlphaEval评估完整的代理产品(如Claude Code、Codex等)作为商业系统,揭示了模型层面评估无法捕捉的性能差异。评估框架融合多种范式(大模型评标、参考驱动度量、形式化验证、评分表评估、自动化UI测试等),各领域组合使用多种评估方式。除基准外,我们还提出一种从真实需求到评估任务的构建框架,实现从需求到评估的标准化、模块化、可复现流程,任何组织均可据此为自身领域构建生产级评估基准。
原文摘要 · Abstract (English)
The rapid deployment of AI agents in commercial settings has outpaced the development of evaluation methodologies that reflect production realities. Existing benchmarks measure agent capabilities through retrospectively curated tasks with well-specified requirements and deterministic metrics -- conditions that diverge fundamentally from production environments where requirements contain implicit constraints, inputs are heterogeneous multi-modal documents with information fragmented across sources, tasks demand undeclared domain expertise, outputs are long-horizon professional deliverables, and success is judged by domain experts whose standards evolve over time. We present AlphaEval, a production-grounded benchmark of 94 tasks sourced from seven companies deploying AI agents in their core business, spanning six O*NET (Occupational Information Network) domains. Unlike model-centric benchmarks, AlphaEval evaluates complete agent products -- Claude Code, Codex, etc. -- as commercial systems, capturing performance variations invisible to model-level evaluation. Our evaluation framework covers multiple paradigms (LLM-as-a-Judge, reference-driven metrics, formal verification, rubric-based assessment, automated UI testing, etc.), with individual domains composing multiple paradigms. Beyond the benchmark itself, we contribute a requirement-to-benchmark construction framework -- a systematic methodology that transforms authentic production requirements into executable evaluation tasks in minimal time. This framework standardizes the entire pipeline from requirement to evaluation, providing a reproducible, modular process that any organization can adopt to construct production-grounded benchmarks for their own domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。