arXiv:2510.18941cs.CLcs.AI2025-10被引 18

构建专业领域评测基准,评估大模型在真实职业场景下的表现

ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge

  • 设计7000+跨领域的专家级评分标准,覆盖物理、化学、金融等专业
  • 顶尖模型在该基准上仅达65.9%准确率,凸显真实专业任务的挑战性
  • 提供低成本可复现的评估方法,适合研究者和开发者验证模型能力

评估大语言模型(LLMs)进展常受限于答案验证难题,现有评测多集中于数学、编程和简短问答。然而,许多真实应用场景需评估模型处理专业文档、整合信息并生成综合报告的能力。我们提出ProfBench:一个包含超过7000个由具备物理学博士、化学博士、金融工商管理硕士及咨询工商管理硕士背景的人类专家评定的响应-标准配对数据集。通过缓解自我增强偏差并降低评估成本2-3个数量级,我们构建了稳健且经济高效的LLM-Judge系统,使评测更公平、普惠。实验发现,即使最先进的模型如GPT-5-high,在ProfBench上也仅取得65.9%的整体性能。此外,我们观察到专有模型与开源模型间存在显著性能差异,并揭示了扩展推理在应对复杂专业任务中的关键作用。数据:https://huggingface.co/datasets/nvidia/ProfBench;代码:https://github.com/NVlabs/ProfBench;排行榜:https://huggingface.co/spaces/nvidia/ProfBench

原文摘要 · Abstract (English)

Evaluating progress in large language models (LLMs) is often constrained by the challenge of verifying responses, limiting assessments to tasks like mathematics, programming, and short-form question-answering. However, many real-world applications require evaluating LLMs in processing professional documents, synthesizing information, and generating comprehensive reports in response to user queries. We introduce ProfBench: a set of over 7000 response-criterion pairs as evaluated by human-experts with professional knowledge across Physics PhD, Chemistry PhD, Finance MBA and Consulting MBA. We build robust and affordable LLM-Judges to evaluate ProfBench rubrics, by mitigating self-enhancement bias and reducing the cost of evaluation by 2-3 orders of magnitude, to make it fair and accessible to the broader community. Our findings reveal that ProfBench poses significant challenges even for state-of-the-art LLMs, with top-performing models like GPT-5-high achieving only 65.9% overall performance. Furthermore, we identify notable performance disparities between proprietary and open-weight models and provide insights into the role that extended thinking plays in addressing complex, professional-domain tasks. Data: https://huggingface.co/datasets/nvidia/ProfBench and Code: https://github.com/NVlabs/ProfBench and Leaderboard: https://huggingface.co/spaces/nvidia/ProfBench

大模型评测专业领域LLM-Judge基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。