arXiv:2503.08404cs.CLcs.CY2025-03被引 5

测试五大模型在政治信息真伪判断中的表现,发现效果有限且存在偏见。

Fact-checking with Generative AI: A Systematic Cross-Topic Examination of LLMs Capacity to Detect Veracity of Political Information

  • 用1.6万条经记者核实的陈述,系统评估5个大模型的查证能力。
  • 各模型准确率不高,对虚假信息识别较好,但敏感话题上差异明显。
  • 适合关注AI查证局限性的研究人员与政策制定者参考。

本研究旨在评估大型语言模型(LLMs)在事实核查中的应用潜力,并推动关于自动化真伪识别的讨论。我们采用AI审计方法,系统评估了五种LLMs(ChatGPT 4、Llama 3(70B)、Llama 3.1(405B)、Claude 3.5 Sonnet、Google Gemini)在16,513条由专业记者验证过的陈述上的表现。通过主题建模与回归分析,探究话题类型、模型种类等因素对真实、虚假及混合类陈述判断的影响。结果表明,尽管ChatGPT 4和Google Gemini表现优于其他模型,整体准确率仍偏低。模型在识别虚假信息方面表现更好,尤其在新冠、美国政治争议和社会议题等敏感领域,提示可能存在影响准确性的潜在约束机制。研究指出,使用LLMs进行事实核查面临显著挑战,包括模型间性能差异大、特定话题输出质量不均,根源或在于训练数据的不足。本文揭示了LLMs在政治事实核查中的潜力与局限,提出未来可通过增强约束机制与微调提升性能。

原文摘要 · Abstract (English)

The purpose of this study is to assess how large language models (LLMs) can be used for fact-checking and contribute to the broader debate on the use of automated means for veracity identification. To achieve this purpose, we use AI auditing methodology that systematically evaluates performance of five LLMs (ChatGPT 4, Llama 3 (70B), Llama 3.1 (405B), Claude 3.5 Sonnet, and Google Gemini) using prompts regarding a large set of statements fact-checked by professional journalists (16,513). Specifically, we use topic modeling and regression analysis to investigate which factors (e.g. topic of the prompt or the LLM type) affect evaluations of true, false, and mixed statements. Our findings reveal that while ChatGPT 4 and Google Gemini achieved higher accuracy than other models, overall performance across models remains modest. Notably, the results indicate that models are better at identifying false statements, especially on sensitive topics such as COVID-19, American political controversies, and social issues, suggesting possible guardrails that may enhance accuracy on these topics. The major implication of our findings is that there are significant challenges for using LLMs for factchecking, including significant variation in performance across different LLMs and unequal quality of outputs for specific topics which can be attributed to deficits of training data. Our research highlights the potential and limitations of LLMs in political fact-checking, suggesting potential avenues for further improvements in guardrails as well as fine-tuning.

事实核查大模型政治信息生成式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。