arXiv:2604.00706cs.CL2026-04

构建非洲语言事实核查数据集,推动低资源语言信息可信度研究

AfrIFact: Cultural Information Retrieval, Evidence Extraction and Fact Checking for African Languages

论文配图:AfrIFact: Cultural Information Retrieval, Evidence Extraction and Fact Checking for African Languages
图 1 · 摘自论文原文
  • 构建涵盖检索、证据提取与验证的非洲语种事实核查数据集
  • 主流模型跨语言检索能力弱,医疗类文档更难检索
  • 少样本提示和微调可显著提升大模型在非洲语言中的事实核查性能

在线言论的真实性评估具有重要现实意义,尤其在信息获取受限、涉及医疗与文化议题的低资源语言社区中更为关键。本文提出AfrIFact数据集,覆盖十种非洲语言及英语的事实核查全流程(信息检索、证据提取、事实验证)。评估显示,即使最佳嵌入模型也缺乏跨语言检索能力,文化与新闻类文档比医疗类文档更易检索,无论在大规模语料或单文档中均如此。大型语言模型在非洲语言中缺乏稳健的多语言事实验证能力,但少样本提示可使AfriqueQwen-14B性能提升达43%,任务特定微调进一步将准确率提高26%。该研究结合数据集发布,推动低资源环境下的信息检索、证据抽取与事实核查研究。

原文摘要 · Abstract (English)

Assessing the veracity of a claim made online is a complex and important task with real-world implications. When these claims are directed at communities with limited access to information and the content concerns issues such as healthcare and culture, the consequences intensify, especially in low-resource languages. In this work, we introduce AfrIFact, a dataset that covers the necessary steps for automatic fact-checking (i.e., information retrieval, evidence extraction, and fact checking), in ten African languages and English. Our evaluation results show that even the best embedding models lack cross-lingual retrieval capabilities, and that cultural and news documents are easier to retrieve than healthcare-domain documents, both in large corpora and in single documents. We show that LLMs lack robust multilingual fact-verification capabilities in African languages, while few-shot prompting improves performance by up to 43% in AfriqueQwen-14B, and task-specific fine-tuning further improves fact-checking accuracy by up to 26%. These findings, along with our release of the AfrIFact dataset, encourage work on low-resource information retrieval, evidence retrieval, and fact checking.

事实核查非洲语言低资源语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。