arXiv:2607.23123cs.AI2026-07

评估大模型在真实工作流中交付任务的能力,发现完成度高不等于质量好。

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

论文配图:SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
图 1 · 摘自论文原文
  • 设计标准化任务链,从基础能力到业务场景分层评估
  • 60.5%最高通过率,三层任务中第三层严格通过率仅18.5%
  • 强调交付风险需独立报告,4.8%的成果因格式等问题未达标

现有大模型评估多集中于知识、推理、编程和工具使用,但极少将受限工作流中生成可验证成果作为评价单位。我们提出SQBench,一个面向生产环境任务交付的语言模型智能体评估基准。SQBench v1.0包含220个标准化任务,分为L1原子能力、L2复合技能和L3业务场景。每个任务要求智能体处理输入资产、调用可用工具,并产出明确指定的交付物。评估先计算功能完成度(Completion),再基于10维风险矩阵中的独立证据推导风险惩罚(Risk Penalty)与性能。严格通过需满足Completion=1且风险惩罚=0。我们在统一协议下评估27种模型配置,每配置-任务对运行一次。最高预设加权通过率@1为60.5%。L3任务平均严格通过率@1为18.5%,所有配置在L3表现均弱于L1和L2,表明在领域约束下交付是当前共性短板。2,348次完成度为1的结果中,113次(4.8%)因不可验证引用、资源误用或格式错误等风险未通过严格标准。结果表明,功能完成不能全面反映交付质量,风险判定应单独报告。

原文摘要 · Abstract (English)

Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced within a constrained workflow as the unit of evaluation. We introduce SQBench, a benchmark for evaluating production-oriented task delivery by language-model agents. SQBench v1.0 contains 220 standardized tasks organized into L1 atomic capabilities, L2 composite skills, and L3 business scenarios. Each task requires an agent to process input assets, use available tools, and produce an explicitly specified deliverable. The evaluation first computes functional Completion and then derives Risk Penalty and Performance from independently evidenced triggers in a 10D Risk Matrix. A Strict Pass requires Completion = 1 and Risk Penalty = 0. We evaluate 27 model configurations under a common protocol, with one run per configuration-task pair. The highest prespecified Weighted Pass@1 is 60.5%. Mean Strict Pass@1 on L3 is 18.5%, and every configuration performs worse on L3 than on both L1 and L2, indicating that delivery under domain constraints is a shared weakness within the current task set. Of 2,348 results with Completion = 1, 113 (4.8%) fail the Strict Pass criterion because of risks such as unverifiable citations, inappropriate resource use, or format violations. These results show that functional completion alone does not fully characterize delivery quality and that risk determinations should be reported separately.

智能体评估任务交付风险检测生产应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。