构建更可靠的评测集,精准衡量大模型的问答真实性。
SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
- 通过多阶段过滤消除标签噪声和题目重复,提升评测可信度。
- 新基准上Gemini 2.5 Pro取得55.6的F1分数,领先GPT-5等模型。
- 适合关注模型幻觉、事实性评估的研究者与开发者使用。
我们提出SimpleQA Verified,一个基于OpenAI SimpleQA的1000个提示的短形式事实性评测基准,用于评估大语言模型(LLM)的事实准确性。该基准解决了原版存在的标签噪声、主题偏差和问题重复等关键缺陷,通过去重、主题均衡和来源校验的多阶段严格过滤流程,构建出更可靠且更具挑战性的评估集,并优化了自动评分提示。在该新基准上,Gemini 2.5 Pro达到55.6的F1分数,超越其他前沿模型(包括GPT-5)。本工作为研究社区提供了一个更高保真度的工具,可真实追踪参数化模型在事实性方面的进展,有效缓解幻觉问题。评测数据集、代码及排行榜已公开:https://www.kaggle.com/benchmarks/deepmind/simpleqa-verified。
原文摘要 · Abstract (English)
We introduce SimpleQA Verified, a 1,000-prompt benchmark for evaluating Large Language Model (LLM) short-form factuality based on OpenAI's SimpleQA. It addresses critical limitations in OpenAI's benchmark, including noisy and incorrect labels, topical biases, and question redundancy. SimpleQA Verified was created through a rigorous multi-stage filtering process involving de-duplication, topic balancing, and source reconciliation to produce a more reliable and challenging evaluation set, alongside improvements in the autorater prompt. On this new benchmark, Gemini 2.5 Pro achieves a state-of-the-art F1-score of 55.6, outperforming other frontier models, including GPT-5. This work provides the research community with a higher-fidelity tool to track genuine progress in parametric model factuality and to mitigate hallucinations. The benchmark dataset, evaluation code, and leaderboard are available at: https://www.kaggle.com/benchmarks/deepmind/simpleqa-verified.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。