利用语言模型的推理能力提升事实核查准确率
Entailed Opinion Matters: Improving the Fact-Checking Performance of Language Models by Relying on their Entailment Ability
- 用生成式模型的推理结果训练编码器模型
- 在多个数据集上达到更高事实核查准确率
- 适合需要可解释性事实核查的应用场景
自动化事实核查对研究社区仍是挑战。以往工作尝试了端到端训练、检索增强生成和提示工程等策略构建鲁棒系统,但准确率仍不足以投入实际应用。本文提出一种新学习范式:利用生成式语言模型(GLMs)生成的证据分类与蕴含推理,来训练编码器语言模型(ELMs)。我们进行了严谨实验,对比了最新方法及多种提示与微调策略,并开展消融实验、错误分析、模型解释质量分析与领域泛化研究,全面验证了该方法的有效性。
原文摘要 · Abstract (English)
Automated fact-checking has been a challenging task for the research community. Prior work has explored various strategies, such as end-to-end training, retrieval-augmented generation, and prompt engineering, to build robust fact-checking systems. However, their accuracy has not been high enough for real-world deployment. We, on the other hand, propose a new learning paradigm, where evidence classification and entailed justifications made by generative language models (GLMs) are used to train encoder-only language models (ELMs). We conducted a rigorous set of experiments, comparing our approach with recent works along with various prompting and fine-tuning strategies. Additionally, we performed ablation studies, error analysis, quality analysis of model explanations, and a domain generalisation study to provide a comprehensive understanding of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。