arXiv:2512.10791cs.CLcs.AI2025-12被引 14

评测大模型在多种场景下的事实准确性,提供全面可靠的基准。

The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality

  • 构建四大子榜单,覆盖图文问答、知识问答、搜索推理和文档对齐。
  • 采用自动化评分模型,综合四项指标得出整体事实性得分。
  • 支持公开与私有数据集,适合评估模型真实能力与应用效果。

我们推出了The FACTS Leaderboard,一个在线排行榜套件及配套基准测试集,全面评估语言模型在多样化场景中生成事实准确文本的能力。该套件通过四个独立子排行榜综合衡量事实性:(1) FACTS Multimodal,评估基于图像问题的回答准确性;(2) FACTS Parametric,通过内部参数回答封闭式事实类问题来检验世界知识;(3) FACTS Search,评估模型使用搜索API进行信息检索时的事实可靠性;(4) FACTS Grounding (v2),评估长文本回复是否基于所提供文档,采用改进的裁判模型。每个子榜单均使用自动化裁判模型评分,最终得分是四项指标的平均值,旨在提供稳健且平衡的整体事实性评估。该套件将持续维护,包含公共与私有数据划分,支持外部参与并保障评估完整性。可访问 https://www.kaggle.com/benchmarks/google/facts 获取。

原文摘要 · Abstract (English)

We introduce The FACTS Leaderboard, an online leaderboard suite and associated set of benchmarks that comprehensively evaluates the ability of language models to generate factually accurate text across diverse scenarios. The suite provides a holistic measure of factuality by aggregating the performance of models on four distinct sub-leaderboards: (1) FACTS Multimodal, which measures the factuality of responses to image-based questions; (2) FACTS Parametric, which assesses models' world knowledge by answering closed-book factoid questions from internal parameters; (3) FACTS Search, which evaluates factuality in information-seeking scenarios, where the model must use a search API; and (4) FACTS Grounding (v2), which evaluates whether long-form responses are grounded in provided documents, featuring significantly improved judge models. Each sub-leaderboard employs automated judge models to score model responses, and the final suite score is an average of the four components, designed to provide a robust and balanced assessment of a model's overall factuality. The FACTS Leaderboard Suite will be actively maintained, containing both public and private splits to allow for external participation while guarding its integrity. It can be found at https://www.kaggle.com/benchmarks/google/facts .

大模型评测事实性基准测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。