为长周期企业任务设计主观评估框架,提升大模型评价可靠性
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks
- 构建专家制定评分标准、真实成果参考与人工偏好对比三支柱评估体系
- 专家评分可信度(kappa=0.60)显著高于大模型自建标准(kappa=0.46)
- 适用于需要长期多工具协作的智能体评估,如设计转代码、内容生成
大语言模型在数学和编程等客观任务上表现优异,评估可简化为单元测试或单一正确答案。然而,现实企业工作常具主观性和情境依赖性:成功取决于组织目标、用户意图以及跨多步骤流程产生的中间成果质量。我们提出LH-Bench,一种三支柱评估设计,超越二元正确性,对主观企业任务中的自主长周期执行进行评分。三支柱包括:(i) 专家制定的评分标准,为大模型裁判提供领域上下文以评估主观工作;(ii) 精选的真实成果作为分步奖励信号(如内容任务中按章节标注);(iii) 成对的人工偏好评估用于收敛验证。实验表明,领域专家制定的评分标准比大模型自建标准更可靠(kappa = 0.60 vs. 0.46),且人工偏好判断证实了顶尖模型间的显著差异(p < 0.05),证明专家驱动评估可在不牺牲可靠性前提下规模化。我们公开发布数据集,并报告在两个环境上的结果:Figma-to-code(33个真实. fig任务通过MCP调用Figma API)、Programmatic content(41门课程共183个独立评估章节,运行于服务30+日活用户的课程平台)。
原文摘要 · Abstract (English)
Large language models excel on objectively verifiable tasks such as math and programming, where evaluation reduces to unit tests or a single correct answer. In contrast, real-world enterprise work is often subjective and context-dependent: success hinges on organizational goals, user intent, and the quality of intermediate artifacts produced across long, multi-tool workflows. We introduce LH-Bench, a three-pillar evaluation design that moves beyond binary correctness to score autonomous, long-horizon execution on subjective enterprise tasks. The pillars are: (i) expert-grounded rubrics that give LLM judges the domain context needed to score subjective work, (ii) curated ground-truth artifacts that enable stepwise reward signals (e.g., chapter-level annotation for content tasks), and (iii) pairwise human preference evaluation for convergent validation. We show that domain-authored rubrics provide substantially more reliable evaluation signals than LLM-authored rubrics (kappa = 0.60 vs. 0.46), and that human preference judgments confirm the same top-tier separation (p < 0.05), evidence that expert-grounded evaluation can scale without sacrificing reliability. We release public datasets and report results on two environments: Figma-to-code (33 real .fig tasks against the Figma API via MCP) and Programmatic content (41 courses comprising 183 individually-evaluated chapters on a course platform serving 30+ daily users).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。