为测试大模型真实学术能力,推出人类最后的闭合式学术测验。
Humanity's Last Exam
- 构建跨学科、需专家知识的多模态学术测验,杜绝网络检索作弊
- 2500道题覆盖数十领域,顶尖大模型准确率不足90%
- 适合评估模型真实学术水平,助力科研与政策制定
基准测试是追踪大语言模型快速进步的重要工具。然而,现有基准难度未同步提升:当前大模型在MMLU等主流基准上准确率已超90%,限制了对模型前沿能力的精准衡量。为此,我们提出人类最后的考试(Humanity's Last Exam, HLE),一个位于人类知识前沿的多模态基准,旨在成为此类闭合式学术基准的终极版本,涵盖广泛学科。HLE包含2500道题目,覆盖数学、人文与自然科学等多个领域。该基准由全球各领域专家共同开发,题目形式包括选择题和简答题,支持自动化评分。每道题均有明确且可验证的答案,但无法通过快速网络检索获得。当前顶尖大模型在HLE上表现不佳,准确率与校准度均较低,凸显了现有模型能力与专家人类水平之间的显著差距。为推动研究与政策制定建立在对模型能力的清晰认知之上,我们已公开发布HLE,网址为https://lastexam.ai。
原文摘要 · Abstract (English)
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve over 90\% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, we introduce Humanity's Last Exam (HLE), a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. HLE consists of 2,500 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. HLE is developed globally by subject-matter experts and consists of multiple-choice and short-answer questions suitable for automated grading. Each question has a known solution that is unambiguous and easily verifiable, but cannot be quickly answered via internet retrieval. State-of-the-art LLMs demonstrate low accuracy and calibration on HLE, highlighting a significant gap between current LLM capabilities and the expert human frontier on closed-ended academic questions. To inform research and policymaking upon a clear understanding of model capabilities, we publicly release HLE at https://lastexam.ai.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。