评测大模型在专家级长文本生成中的表现,提出可量化评估框架。
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists
- 基于领域专家设计的检查清单,对长文本输出进行结构化评估。
- 现有模型最高仅33.4分F1,生成内容虽覆盖要点但准确性不足。
- 开源评估框架支持低成本、可复现的专家级生成质量评测。
本文提出ExpertLongBench,一个包含9个领域11项任务的专家级基准,涵盖真实工作流程与应用需求。任务要求生成超5000词的长文本,并严格遵守领域规范。每项任务配有由领域专家设计或验证的评分标准。我们提出CLEAR评估框架,通过从模型输出与参考答案中提取对应检查清单项,实现细粒度、专家对齐的评估。实验评估13个主流大模型,结果显示:当前最优模型Gemini-2.5-Pro仅获33.4 F1,表明模型虽能覆盖所需方面,但准确性严重不足;同时,开放权重模型可实现准确的清单提取与比对,支持更高效、可复现、低成本的评估。
原文摘要 · Abstract (English)
This paper introduces ExpertLongBench, an expert-level benchmark containing 11 tasks from 9 domains that reflect realistic expert workflows and applications. Beyond question answering, the application-driven tasks in ExpertLongBench demand long-form outputs that can exceed 5,000 tokens and strict adherence to domain-specific requirements. Notably, each task in ExpertLongBench includes a rubric, designed or validated by domain experts, to specify task requirements and guide output evaluation. Furthermore, we propose CLEAR, an evaluation framework that supports accurate evaluation of long-form model outputs in our benchmark. To achieve fine-grained, expert-aligned evaluation, CLEAR derives checklists from both model outputs and references by extracting information corresponding to items in the task-specific rubric. Checklist items of model outputs are then compared with corresponding items of reference outputs to assess their correctness, enabling grounded evaluation. We benchmark 13 popular large language models (LLMs) and analyze components in CLEAR, showing that (1) existing LLMs, with the top performer Gemini-2.5-Pro achieving only a 33.4 F1 score, require significant improvement for expert-level tasks; (2) models can generate content corresponding to the required aspects, but far from correct; and (3) accurate checklist extraction and comparison in CLEAR can be achieved by open-weight models for more scalable, reproducible, and low-cost usage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。