修复了人类终极测试中的错误数据,提升大模型评估可信度
HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
- 通过专家评审与模型交叉验证,筛选出668个可靠题目
- 对可修正题目进行双人修订与模型审计,新增1143个认证题
- 发现模型自信度与题目错误高度相关,适合评估可信性研究
人类终极测试(HLE)是评估前沿大语言模型在复杂多领域问题上的主流基准,但社区分析指出其包含大量噪声数据,可能扭曲评估结果与模型对比。为此,我们提出HLE-Verified,一个经过透明验证与精细修订的改进版本。第一阶段通过领域专家评审和模型交叉验证,完成668项题目的二元验证;第二阶段对可修正题目实施双重独立专家修订、模型辅助审核及最终裁决,在保持原评估意图前提下生成1143个修订并认证的题目。剩余689项被归为标注不确定集,附有明确不确定性来源与专业标签以供后续优化。在八个先进语言模型上测试发现,相较于原HLE,HLE-Verified平均准确率提升7–10个百分点,尤其在原始问题或参考答案错误的题目上提升达30–40个百分点。分析还揭示模型置信度与题目错误存在强关联,验证了修订有效性。该工作显著降低标注噪声,推动更真实的模型能力测量。数据已公开于:https://huggingface.co/datasets/skylenage/HLE-Verified
原文摘要 · Abstract (English)
Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions. However, community-led analyses have raised concerns that HLE contains a non-trivial number of noisy items, which can bias evaluation results and distort cross-model comparisons. To address this challenge, we introduce HLE-Verified, a verified and revised version of HLE with a transparent verification protocol and fine-grained error taxonomy. Our construction follows a two-stage validation-and-repair workflow resulting in a certified benchmark. In Stage I, each item undergoes binary validation of the problem and final answer through domain-expert review and model-based cross-checks, yielding 668 verified items. In Stage II, flawed but fixable items are revised under strict constraints preserving the original evaluation intent, through dual independent expert repairs, model-assisted auditing, and final adjudication, resulting in 1,143 revised-and-certified items. The remaining 689 items are released as a documented uncertain set with explicit uncertainty sources and expertise tags for future refinement. We evaluate eight state-of-the-art language models on HLE and HLE-Verified, observing an average absolute accuracy gain of 7--10 percentage points on HLE-Verified. The improvement is particularly pronounced on items where the original problem statement and/or reference answer is erroneous, with gains of 30--40 percentage points. Our analyses further reveal a strong association between model confidence and the presence of errors in the problem statement or reference answer, supporting the effectiveness of our revisions. Overall, HLE-Verified improves HLE-style evaluations by reducing annotation noise and enabling more faithful measurement of model capabilities. Data is available at: https://huggingface.co/datasets/skylenage/HLE-Verified
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。