arXiv:2503.17684cs.CLcs.AI2025-03Transactions of th…被引 11

用大模型自动生成可公开传播的新闻核查文章。

Can LLMs Automate Fact-Checking Article Writing?

  • 设计QRAFT框架,模仿人工核查员写作流程。
  • 人类评估显示其生成文章质量低于专家水平。
  • 为自动化核查输出提供新方向,适合研究者参考。

自动核查旨在通过工具辅助专业核查人员,提升人工核查效率。然而,现有框架未能解决关键环节:生成适合大众传播的核查结果。虽然人工核查员会撰写完整的核查文章,但自动化系统通常缺乏充分论据支持。本文旨在填补这一空白,提出在传统自动核查流程中增加自动生成完整核查文章的能力。我们通过与主流核查机构专家访谈,确定了高质量核查文章的关键要求,并开发了基于大模型的代理框架QRAFT,模拟人类核查员的写作流程。最后,通过专业核查员的人工评估验证QRAFT的实际效用。评估结果显示,尽管QRAFT优于多个已有文本生成方法,但仍显著落后于专家撰写的核查文章。本研究希望推动该新兴重要方向的进一步探索。代码已开源:https://github.com/mbzuai-nlp/qraft.git。

原文摘要 · Abstract (English)

Automatic fact-checking aims to support professional fact-checkers by offering tools that can help speed up manual fact-checking. Yet, existing frameworks fail to address the key step of producing output suitable for broader dissemination to the general public: while human fact-checkers communicate their findings through fact-checking articles, automated systems typically produce little or no justification for their assessments. Here, we aim to bridge this gap. In particular, we argue for the need to extend the typical automatic fact-checking pipeline with automatic generation of full fact-checking articles. We first identify key desiderata for such articles through a series of interviews with experts from leading fact-checking organizations. We then develop QRAFT, an LLM-based agentic framework that mimics the writing workflow of human fact-checkers. Finally, we assess the practical usefulness of QRAFT through human evaluations with professional fact-checkers. Our evaluation shows that while QRAFT outperforms several previously proposed text-generation approaches, it lags considerably behind expert-written articles. We hope that our work will enable further research in this new and important direction. The code for our implementation is available at https://github.com/mbzuai-nlp/qraft.git.

自动核查大模型文章生成信息可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。