构建285个学科的超大规模评估基准,检验大模型在冷门领域的知识水平。
SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- 用人类与大模型协同筛选机制,迭代优化题目质量
- 当前最强模型在该基准上最高准确率仅61.82%
- 适合关注模型泛化能力与跨学科评估的研究者
大语言模型在数学、物理和计算机科学等主流学科中表现优异,但人类知识涵盖200多个专业领域,远超现有评测范围。许多细分领域——尤其是轻工业、农业和服务类学科——的大模型能力尚未充分评估。为此,我们提出SuperGPQA,一个覆盖285个学科的研究生级知识与推理能力评估基准。该基准采用新颖的人类-大模型协同过滤机制,通过结合模型回答与专家反馈进行迭代优化,剔除琐碎或模糊问题。实验结果显示,当前顶尖大模型在多样化知识领域仍存在显著提升空间(如推理导向模型DeepSeek-R1在SuperGPQA上最高准确率为61.82%),凸显现有模型能力与通用人工智能之间的巨大差距。此外,我们还分享了管理大规模标注过程的经验,涉及80多位专家与交互式人机协同系统,为未来同类研究提供方法论参考。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable proficiency in mainstream academic disciplines such as mathematics, physics, and computer science. However, human knowledge encompasses over 200 specialized disciplines, far exceeding the scope of existing benchmarks. The capabilities of LLMs in many of these specialized fields-particularly in light industry, agriculture, and service-oriented disciplines-remain inadequately evaluated. To address this gap, we present SuperGPQA, a comprehensive benchmark that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines. Our benchmark employs a novel Human-LLM collaborative filtering mechanism to eliminate trivial or ambiguous questions through iterative refinement based on both LLM responses and expert feedback. Our experimental results reveal significant room for improvement in the performance of current state-of-the-art LLMs across diverse knowledge domains (e.g., the reasoning-focused model DeepSeek-R1 achieved the highest accuracy of 61.82% on SuperGPQA), highlighting the considerable gap between current model capabilities and artificial general intelligence. Additionally, we present comprehensive insights from our management of a large-scale annotation process, involving over 80 expert annotators and an interactive Human-LLM collaborative system, offering valuable methodological guidance for future research initiatives of comparable scope.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。