构建大规模多语言表格事实核查数据集,推动真实世界结构化数据验证研究。
Frame-Guided Synthetic Claim Generation for Automatic Fact-Checking Using High-Volume Tabular Data
- 基于六种语义框架自动生成可信度高的合成声明。
- 每张表超50万行,共78,503条声明,覆盖4种语言。
- 揭示大模型需真正检索而非记忆,适合研发事实核查系统者使用。
自动化事实核查基准长期忽视对高容量结构化数据的验证挑战,多聚焦于小规模精调表格。本文提出一个大规模、多语言数据集,包含78,503条基于434张复杂OECD表格生成的合成声明,每张表平均超过50万行。我们设计了一种新型帧引导方法,算法依据六种语义框架程序化选取关键数据点,生成英语、中文、西班牙语和印地语的合理声明。通过知识探测实验,我们证明大型语言模型未记忆这些事实,迫使系统进行真实的数据检索与推理。我们提供一个基准SQL生成系统,表明该基准极具挑战性。分析显示,证据检索是主要瓶颈,模型难以在海量表格中定位正确数据。该数据集为解决这一现实世界难题提供了关键资源。
原文摘要 · Abstract (English)
Automated fact-checking benchmarks have largely ignored the challenge of verifying claims against real-world, high-volume structured data, instead focusing on small, curated tables. We introduce a new large-scale, multilingual dataset to address this critical gap. It contains 78,503 synthetic claims grounded in 434 complex OECD tables, which average over 500K rows each. We propose a novel, frame-guided methodology where algorithms programmatically select significant data points based on six semantic frames to generate realistic claims in English, Chinese, Spanish, and Hindi. Crucially, we demonstrate through knowledge-probing experiments that LLMs have not memorized these facts, forcing systems to perform genuine retrieval and reasoning rather than relying on parameterized knowledge. We provide a baseline SQL-generation system and show that our benchmark is highly challenging. Our analysis identifies evidence retrieval as the primary bottleneck, with models struggling to find the correct data in massive tables. This dataset provides a critical new resource for advancing research on this unsolved, real-world problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。