用大模型生成多语言事实核查数据,解决低资源语言信息差问题。
Beyond Translation: LLM-Based Data Generation for Multilingual Fact-Checking
- 用LLM结合维基百科生成跨语言事实核查对
- 构建220万条多语言声明-来源配对数据集
- 开源工具链,助力多语言事实核查研究
强大的自动事实核查系统有望大规模遏制网络虚假信息。然而,现有研究主要集中在英语。本文提出MultiSynFact,首个包含220万条声明-来源配对的大规模多语言事实核查数据集,支持西班牙语、德语、英语及其他低资源语言。数据生成流程利用大型语言模型(LLMs),融合维基百科外部知识,并通过严格的声明验证步骤确保质量。我们在多种模型和实验设置下评估了MultiSynFact的有效性。此外,我们开源了一个用户友好的框架,以促进多语言事实核查与数据生成的进一步研究。
原文摘要 · Abstract (English)
Robust automatic fact-checking systems have the potential to combat online misinformation at scale. However, most existing research primarily focuses on English. In this paper, we introduce MultiSynFact, the first large-scale multilingual fact-checking dataset containing 2.2M claim-source pairs designed to support Spanish, German, English, and other low-resource languages. Our dataset generation pipeline leverages Large Language Models (LLMs), integrating external knowledge from Wikipedia and incorporating rigorous claim validation steps to ensure data quality. We evaluate the effectiveness of MultiSynFact across multiple models and experimental settings. Additionally, we open-source a user-friendly framework to facilitate further research in multilingual fact-checking and dataset generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。