arXiv:2502.08909cs.CLcs.AI2025-02被引 11

用大模型自动核查真实言论,提升准确率与解释力。

Towards Automated Fact-Checking of Real-World Claims: Exploring Task Formulation and Assessment with LLMs

  • 构建多标签体系,结合检索证据实现结构化判断
  • 70B大模型在无微调下表现最优,比小模型更优
  • 适合研究自动化辟谣与模型可信解释的开发者

为应对虚假信息泛滥,传统人工核查效率低、成本高。本研究利用大语言模型(LLMs)开展自动化事实核查(AFC),在二分类、三分类、五分类等多种标注方案下建立基准。基于2007–2024年来自PolitiFact的17,856条真实言论,通过受限网络搜索获取证据,评估Llama-3系列模型(3B、8B、70B)性能。采用无需参考的TIGERScore评分机制衡量解释质量。结果表明:未微调的大模型在分类准确率和解释质量上均优于小模型;3B模型在单次提示下表现接近经过微调的大上下文小模型,但70B模型始终领先。证据融合显著提升所有模型表现,尤其对大模型增益更大。细微标签区分仍具挑战,凸显需进一步优化标注体系与证据对齐。研究验证了检索增强型大模型在真实世界事实核查中的潜力。

原文摘要 · Abstract (English)

Fact-checking is necessary to address the increasing volume of misinformation. Traditional fact-checking relies on manual analysis to verify claims, but it is slow and resource-intensive. This study establishes baseline comparisons for Automated Fact-Checking (AFC) using Large Language Models (LLMs) across multiple labeling schemes (binary, three-class, five-class) and extends traditional claim verification by incorporating analysis, verdict classification, and explanation in a structured setup to provide comprehensive justifications for real-world claims. We evaluate Llama-3 models of varying sizes (3B, 8B, 70B) on 17,856 claims collected from PolitiFact (2007-2024) using evidence retrieved via restricted web searches. We utilize TIGERScore as a reference-free evaluation metric to score the justifications. Our results show that larger LLMs consistently outperform smaller LLMs in classification accuracy and justification quality without fine-tuning. We find that smaller LLMs in a one-shot scenario provide comparable task performance to fine-tuned Small Language Models (SLMs) with large context sizes, while larger LLMs consistently surpass them. Evidence integration improves performance across all models, with larger LLMs benefiting most. Distinguishing between nuanced labels remains challenging, emphasizing the need for further exploration of labeling schemes and alignment with evidences. Our findings demonstrate the potential of retrieval-augmented AFC with LLMs.

自动核查大模型事实验证证据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。