arXiv:2506.03655cs.CYcs.CL2025-06被引 3

多语言事实核查测试揭示大模型在真实场景中的可靠性短板。

Facts are Harder Than Opinions -- A Multilingual, Comparative Analysis of LLM-Based Fact-Checking Reliability

  • 构建包含6万+跨语言跨主题声明的动态扩展数据集
  • GPT-4o准确率最高,但无法处理43%的声明
  • 事实类陈述比观点类更易被误判,提示系统潜在风险

虚假信息泛滥亟需可扩展的自动化事实核查方案。然而现有基准常忽视多语言与主题多样性。本文提出一个新型动态可扩展数据集,涵盖61,514条多语言、多主题的声明,覆盖至2024年。通过对GPT-4o、GPT-3.5 Turbo、LLaMA 3.1和Mixtral 8x7B等五款主流大模型的全面评估,发现不同语言与主题间存在显著性能差异。尽管整体上GPT-4o表现最佳,仍无法分类43%的声明。所有模型均更易将事实性陈述误判为正确,而观点类则相对稳定,揭示出关键脆弱性。研究警示应谨慎部署基于大模型的事实核查系统,并指出其规模化应用的挑战。

原文摘要 · Abstract (English)

The proliferation of misinformation necessitates scalable, automated fact-checking solutions. Yet, current benchmarks often overlook multilingual and topical diversity. This paper introduces a novel, dynamically extensible data set that includes 61,514 claims in multiple languages and topics, extending existing datasets up to 2024. Through a comprehensive evaluation of five prominent Large Language Models (LLMs), including GPT-4o, GPT-3.5 Turbo, LLaMA 3.1, and Mixtral 8x7B, we identify significant performance gaps between different languages and topics. While overall GPT-4o achieves the highest accuracy, it declines to classify 43% of claims. Across all models, factual-sounding claims are misclassified more often than opinions, revealing a key vulnerability. These findings underscore the need for caution and highlight challenges in deploying LLM-based fact-checking systems at scale.

事实核查大模型评测多语言可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。