arXiv:2511.11562cs.CLcs.CY2025-11被引 22

构建首个大规模专业推理评估基准,专为金融法律领域真实难题设计

PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning

  • 用182名专业人士设计1100个真实工作场景任务
  • 顶级模型在核心测试集上得分仅0.39(金融)和0.37(法律)
  • 突出模型推理不透明、判断不准等关键缺陷,适合风控与合规研究者

前沿模型进展常以学术基准衡量,但此类评估难以反映真实职业场景下的表现。现有方法无法有效评测金融与法律等高风险领域中开放性、经济影响重大的任务。为此,我们推出专业推理基准(PRBench),涵盖金融与法律领域的现实问题。该基准开源了1100个由专业人士撰写的任务及19,356条专家评审标准,是目前规模最大的公开、基于评分标准的双领域基准。我们招募了182名持证律师(JD)、注册金融分析师(CFA)或具备6年以上经验的专业人士,其任务灵感源自实际工作流程,覆盖114个国家和47个美国司法管辖区。评审标准经独立专家验证,确保质量。对20个主流模型的评估显示,其在困难子集上的最高分仅为0.39(金融)和0.37(法律),表明仍有巨大提升空间。我们进一步分析任务的经济影响,并基于人工标注的评分类别进行性能拆解,发现相似总分的模型在具体能力上差异显著。常见失败模式包括判断错误、推理过程不透明和逻辑不完整,暴露出模型在专业应用中的可靠性短板。

原文摘要 · Abstract (English)

Frontier model progress is often measured by academic benchmarks, which offer a limited view of performance in real-world professional contexts. Existing evaluations often fail to assess open-ended, economically consequential tasks in high-stakes domains like Legal and Finance, where practical returns are paramount. To address this, we introduce Professional Reasoning Bench (PRBench), a realistic, open-ended, and difficult benchmark of real-world problems in Finance and Law. We open-source its 1,100 expert-authored tasks and 19,356 expert-curated criteria, making it, to our knowledge, the largest public, rubric-based benchmark for both legal and finance domains. We recruit 182 qualified professionals, holding JDs, CFAs, or 6+ years of experience, who contributed tasks inspired by their actual workflows. This process yields significant diversity, with tasks spanning 114 countries and 47 US jurisdictions. Our expert-curated rubrics are validated through a rigorous quality pipeline, including independent expert validation. Subsequent evaluation of 20 leading models reveals substantial room for improvement, with top scores of only 0.39 (Finance) and 0.37 (Legal) on our Hard subsets. We further catalog associated economic impacts of the prompts and analyze performance using human-annotated rubric categories. Our analysis shows that models with similar overall scores can diverge significantly on specific capabilities. Common failure modes include inaccurate judgments, a lack of process transparency and incomplete reasoning, highlighting critical gaps in their reliability for professional adoption.

专业推理金融法律评估基准专家评审

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。