用高质量摘要提升大模型政治事实核查准确率
Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search
- 用权威摘要构建检索增强系统,替代直接调用网络搜索
- 相比标准模型,综合准确率提升233%(平均)
- 适合需要高可靠性的事实核查场景
大语言模型(LLMs)曾被寄予自动化端到端事实核查的厚望,但已有研究结果参差不齐。随着主流聊天机器人陆续集成推理和网络搜索功能,数百万用户已依赖它们进行信息验证,因此亟需严谨评估。我们对来自OpenAI、Google、Meta和DeepSeek的15个最新大模型在超过6,000条由PolitiFact核实的政治声明上进行了评测,对比了标准模型、带推理能力的变体及启用网络搜索的版本。结果显示:标准模型表现不佳,推理带来微弱改进,网络搜索仅提供中等增益,即便这些事实本身已在网页上可查。相反,使用PolitiFact摘要构建的定制化RAG系统,在所有模型变体上平均将宏观F1值提升了233%。这表明,为模型提供经过筛选的高质量上下文是实现可靠自动化事实核查的有效路径。
原文摘要 · Abstract (English)
Large language models (LLMs) have raised hopes for automated end-to-end fact-checking, but prior studies report mixed results. As mainstream chatbots increasingly ship with reasoning capabilities and web search tools -- and millions of users already rely on them for verification -- rigorous evaluation is urgent. We evaluate 15 recent LLMs from OpenAI, Google, Meta, and DeepSeek on more than 6,000 claims fact-checked by PolitiFact, comparing standard models with reasoning- and web-search variants. Standard models perform poorly, reasoning offers minimal benefits, and web search provides only moderate gains, despite fact-checks being available on the web. In contrast, a curated RAG system using PolitiFact summaries improved macro F1 by 233% on average across model variants. These findings suggest that giving models access to curated high-quality context is a promising path for automated fact-checking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。