arXiv:2503.22877cs.CLcs.AI2025-03被引 2

LLM事实核查在不同地区表现差异大,南半球内容更易出错。

Understanding Inequality of LLM Fact-Checking over Geographic Regions with Agent and Retrieval models

  • 用三种场景测试LLM查证能力:仅凭陈述、调用维基百科的智能体、最佳情况下的检索增强系统
  • 无论模型或场景,全球北方的陈述准确率显著高于全球南方,差距在维基智能体场景中扩大
  • 提示需关注地理不平衡问题,适合研究AI公平性与信息核查的学者参考

事实核查是大型语言模型(LLMs)对抗虚假信息传播的潜在应用。然而,LLM在不同地理区域的表现存在差异。本文评估了开源与私有模型在多种区域和情境下的事实准确性。基于包含600条经核实陈述的平衡数据集,覆盖六个全球区域,我们考察了三种事实核查场景:(1) 仅提供陈述;(2) 使用具备维基百科访问权限的LLM智能体;(3) 在提供官方核查结果的前提下,采用检索增强生成(RAG)系统作为最优情况。结果表明,无论场景或所用模型(包括GPT-4、Claude Sonnet、LLaMA),全球北方的陈述表现均显著优于全球南方。该差距在维基百科智能体系统中进一步扩大,凸显通用知识库难以应对地域性细节。研究强调亟需改进数据集平衡与鲁棒检索策略,以提升多地理语境下LLM的事实核查能力。

原文摘要 · Abstract (English)

Fact-checking is a potentially useful application of Large Language Models (LLMs) to combat the growing dissemination of disinformation. However, the performance of LLMs varies across geographic regions. In this paper, we evaluate the factual accuracy of open and private models across a diverse set of regions and scenarios. Using a dataset containing 600 fact-checked statements balanced across six global regions we examine three experimental setups of fact-checking a statement: (1) when just the statement is available, (2) when an LLM-based agent with Wikipedia access is utilized, and (3) as a best case scenario when a Retrieval-Augmented Generation (RAG) system provided with the official fact check is employed. Our findings reveal that regardless of the scenario and LLM used, including GPT-4, Claude Sonnet, and LLaMA, statements from the Global North perform substantially better than those from the Global South. Furthermore, this gap is broadened for the more realistic case of a Wikipedia agent-based system, highlighting that overly general knowledge bases have a limited ability to address region-specific nuances. These results underscore the urgent need for better dataset balancing and robust retrieval strategies to enhance LLM fact-checking capabilities, particularly in geographically diverse contexts.

LLM事实核查地理偏见RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。