arXiv:2503.05565cs.CYcs.CL2025-03被引 7

评测开源大模型在事实核查中的表现,发现其有潜力但仍有明显短板。

Evaluating open-source Large Language Models for automated fact-checking

  • 通过三类实验测试大模型识别论断与核查文章关系的能力。
  • 在已有核查文章时准确率较高,但对新事实判断能力弱于传统模型。
  • 引入外部知识未能显著提升性能,需更针对性优化方法。

网络虚假信息泛滥加剧了对自动化事实核查的需求。大语言模型(LLMs)被视为潜在工具,但其有效性尚不明确。本研究评估多种开源LLM在不同上下文信息条件下的事实核查能力。开展三项关键实验:(1) 判断论断与核查文章间的语义关系;(2) 在提供相关核查文章时验证论断的准确性;(3) 利用Google和Wikipedia等外部知识源进行核查。结果表明,LLMs在识别论断-文章关联及验证已核查内容方面表现良好,但在确认新事实时表现不佳,不及如RoBERTa等传统微调模型。引入外部知识未显著提升性能,提示需发展更适配的方法。研究揭示了LLMs在事实核查中的潜力与局限,强调其尚无法可靠替代人工核查,需进一步优化。

原文摘要 · Abstract (English)

The increasing prevalence of online misinformation has heightened the demand for automated fact-checking solutions. Large Language Models (LLMs) have emerged as potential tools for assisting in this task, but their effectiveness remains uncertain. This study evaluates the fact-checking capabilities of various open-source LLMs, focusing on their ability to assess claims with different levels of contextual information. We conduct three key experiments: (1) evaluating whether LLMs can identify the semantic relationship between a claim and a fact-checking article, (2) assessing models' accuracy in verifying claims when given a related fact-checking article, and (3) testing LLMs' fact-checking abilities when leveraging data from external knowledge sources such as Google and Wikipedia. Our results indicate that LLMs perform well in identifying claim-article connections and verifying fact-checked stories but struggle with confirming factual news, where they are outperformed by traditional fine-tuned models such as RoBERTa. Additionally, the introduction of external knowledge does not significantly enhance LLMs' performance, calling for more tailored approaches. Our findings highlight both the potential and limitations of LLMs in automated fact-checking, emphasizing the need for further refinements before they can reliably replace human fact-checkers.

事实核查大模型开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。