arXiv:2505.21409cs.CLcs.AI2025-05

测试大模型从表格中提取结构化事实的能力,发现其准确率不足25%。

RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models

  • 设计新基准,评估模型生成多行多列表格的准确率
  • 最先进模型在复杂查询下准确率低于25%,且随输出规模下降
  • 适合关注大模型事实性与结构化输出的研究者使用

大语言模型的事实性仍是长期挑战。现有评测多聚焦短答案,忽视了从参数化知识中生成结构化、多记录表格的能力。我们证明,即使单个事实已知,关系型事实检索比单一查询困难得多,且失败模式与输出维度(如属性数或记录数)密切相关。为系统评估这一被低估的能力,我们提出RelationalFactQA,包含多样自然语言问题(配对应SQL)和标准表格答案,专用于评估结构化知识检索。该基准支持分析不同查询复杂度、输出规模和数据特征下的表现。实验显示,即使最先进的模型在生成关系输出时事实准确率也未超过25%,且随着输出维度增加显著下降。这些结果凸显当前大模型在合成结构化事实知识方面的严重局限,并确立RelationalFactQA作为未来衡量大模型事实性进展的关键资源。

原文摘要 · Abstract (English)

Factuality in Large Language Models (LLMs) is a persistent challenge. Current benchmarks often assess short factual answers, overlooking the critical ability to generate structured, multi-record tabular outputs from parametric knowledge. We demonstrate that this relational fact retrieval is substantially more difficult than isolated point-wise queries, even when individual facts are known to the model, exposing distinct failure modes sensitive to output dimensionality (e.g., number of attributes or records). To systematically evaluate this under-explored capability, we introduce RelationalFactQA, a new benchmark featuring diverse natural language questions (paired with SQL) and gold-standard tabular answers, specifically designed to assess knowledge retrieval in a structured format. RelationalFactQA enables analysis across varying query complexities, output sizes, and data characteristics. Our experiments reveal that even state-of-the-art LLMs struggle significantly, not exceeding 25% factual accuracy in generating relational outputs, with performance notably degrading as output dimensionality increases. These findings underscore critical limitations in current LLMs' ability to synthesize structured factual knowledge and establish RelationalFactQA as a crucial resource for measuring future progress in LLM factuality.

大模型事实性表格生成评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。