构建首个经人工验证的长文本事实性评测集,揭示大模型事实错误率高达40%。
FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality
- 采用人机协同生成挑战性事实类问题,确保问题可答且无歧义。
- 顶尖模型在该数据集上事实错误率达40%,远超其他数据集的10%。
- 适合评估模型跨长尾知识推理能力,推动可信AI发展。
长文本事实性评估旨在衡量模型对简短提示生成准确、全面回答的能力。现有基准普遍缺乏人工验证,存在质量隐患。为解决这一问题,我们提出FACTORY,一个大规模、经人工验证的提示集合。通过模型在环方法生成并由人类优化,FACTORY包含具有挑战性的事实型问题,要求答案可获取、明确且无需猜测。我们在6个前沿语言模型上使用FACTORY及现有数据集进行人工评估。结果表明,FACTORY是一个极具挑战性的基准:顶尖模型生成内容中约40%的陈述不属实,而其他数据集仅为10%。分析显示,FACTORY相比以往基准更具可靠性,凸显模型需具备处理长尾事实的推理能力。
原文摘要 · Abstract (English)
Long-form factuality evaluation assesses the ability of models to generate accurate, comprehensive responses to short prompts. Existing benchmarks often lack human verification, leading to potential quality issues. To address this limitation, we introduce FACTORY, a large-scale, human-verified prompt set. Developed using a model-in-the-loop approach and refined by humans, FACTORY includes challenging prompts that are fact-seeking, answerable, and unambiguous. We conduct human evaluations on 6 state-of-the-art language models using FACTORY and existing datasets. Our results show that FACTORY is a challenging benchmark: approximately 40% of the claims made in the responses of SOTA models are not factual, compared to only 10% for other datasets. Our analysis identifies the strengths of FACTORY over prior benchmarks, emphasizing its reliability and the necessity for models to reason across long-tailed facts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。