构建专家级任务评估基准,突破大模型在真实专业场景中的能力瓶颈。
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation

- 设计80个领域1346项专家任务,覆盖金融、医疗等真实职业场景。
- 采用15-40个加权评分点的详细评分标准,确保评估专业严谨性。
- 提出新评估范式ShotJudge,有效减少模型自评偏差,适合专业级模型评测。
随着大语言模型在传统基准上表现趋于饱和,一个核心挑战仍存:如何评估其在复杂、开放式任务中体现的真实专家级认知能力。现有框架普遍存在领域覆盖窄、依赖通用任务或自评偏差等问题。为此,我们提出XpertBench,一个高保真度的基准,用于评估大模型在真实专业领域的表现。XpertBench包含80个类别共1346项精心设计的任务,涵盖金融、医疗、法律、教育及理工与人文双轨研究。任务源自超过1000份领域专家(包括顶尖机构研究员和具有丰富临床/工业经验的从业者)的提交,确保生态有效性。每项任务配备15-40个加权评分点的详细评分细则,以衡量专业严谨性。为实现可扩展且符合人类判断的评估,我们引入ShotJudge——一种使用专家少样本示例校准的大模型裁判机制,以缓解自奖励偏差。对主流大模型的实证评估显示显著性能天花板:即使顶尖模型最高成功率达约66%,平均分仅约55%。模型还表现出领域特异性差异,在定量推理与语言合成方面呈现非重叠优势。这些发现凸显当前AI系统存在显著的“专家鸿沟”,并确立XpertBench作为推动通用助手向专业化协作伙伴演进的关键工具。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing frameworks suffer from narrow domain coverage, reliance on generalist tasks, or self-evaluation biases. To bridge this gap, we present XpertBench, a high-fidelity benchmark engineered to assess LLMs across authentic professional domains. XpertBench consists of 1,346 meticulously curated tasks across 80 categories, spanning finance, healthcare, legal services, education, and dual-track research (STEM and Humanities). These tasks are derived from over 1,000 submissions by domain experts--including researchers from elite institutions and practitioners with extensive clinical or industrial experience--ensuring superior ecological validity. Each task uses detailed rubrics with mostly 15-40 weighted checkpoints to assess professional rigor. To facilitate scalable yet human-aligned assessment, we introduce ShotJudge, a novel evaluation paradigm that employs LLM judges calibrated with expert few-shot exemplars to mitigate self-rewarding biases. Our empirical evaluation of state-of-the-art LLMs reveals a pronounced performance ceiling: even leading models achieve a peak success rate of only ~66%, with a mean score around 55%. Models also exhibit domain-specific divergence, showing non-overlapping strengths in quantitative reasoning versus linguistic synthesis.. These findings underscore a significant "expert-gap" in current AI systems and establish XpertBench as a critical instrument for navigating the transition from general-purpose assistants to specialized professional collaborators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。