评测大模型检索增强生成在专业事实核查中的表现
Face the Facts! Evaluating RAG-based Pipelines for Professional Fact-Checking
- 用真实复杂语境和多样可靠知识库测试RAG生成结论的能力
- 大模型更忠实于事实,小模型更贴合上下文,零/一跳提示信息量更高
- 适合研究自动化核查系统或需平衡准确与风格的开发者
自然语言处理与生成系统近期展现出辅助专业事实核查人员、提升效率的潜力。本文突破现有基于检索增强生成(RAG)的自动核查流水线的局限,遵循专业核查实践,对生成结论(即判断声明真伪的短文本)的RAG方法进行基准测试,评估其在语义复杂的声明和异构但可靠的知识库上的表现。结果表明:基于大模型的检索器优于其他检索技术,但在异构知识库上仍存困难;更大模型在结论忠实度上表现更佳,而更小模型在上下文贴合度上更优;人工评估显示,零样本与单样本方法在信息量上更受青睐,微调模型则在情感一致性上表现更好。
原文摘要 · Abstract (English)
Natural Language Processing and Generation systems have recently shown the potential to complement and streamline the costly and time-consuming job of professional fact-checkers. In this work, we lift several constraints of current state-of-the-art pipelines for automated fact-checking based on the Retrieval-Augmented Generation (RAG) paradigm. Our goal is to benchmark, following professional fact-checking practices, RAG-based methods for the generation of verdicts - i.e., short texts discussing the veracity of a claim - evaluating them on stylistically complex claims and heterogeneous, yet reliable, knowledge bases. Our findings show a complex landscape, where, for example, LLM-based retrievers outperform other retrieval techniques, though they still struggle with heterogeneous knowledge bases; larger models excel in verdict faithfulness, while smaller models provide better context adherence, with human evaluations favouring zero-shot and one-shot approaches for informativeness, and fine-tuned models for emotional alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。