arXiv:2410.12377cs.CLcs.CY2024-10被引 20

用开源大模型集群实现真实世界声明的自动化验证,效果逼近顶尖水平。

HerO at AVeriTeC: The Herd of Open Large Language Models for Verifying Real-World Claims

  • 用大模型生成假设性文档增强查询,提升证据检索精度。
  • 在AVeriTeC任务中取得0.57的得分,排名第二。
  • 全程仅使用公开模型,适合追求可复现性的研究者。

为应对FEVER-24举办的AVeriTeC共享任务,我们提出仅依赖公开可用大语言模型(LLMs)完成自动事实核查全过程的系统,名为「赫尔多:用于验证现实声明的开源大模型集群」(HerO)。在证据检索阶段,通过大模型生成假设性核查文档以增强查询;在问题生成与真伪判断环节,采用包含检索到的上下文样本的提示工程,调用预训练和微调后的LLMs。HerO在排行榜上以0.57的AVeriTeC得分获得第二名,表明开源大模型在真实世界声明验证中的巨大潜力。我们已将代码公开于https://github.com/ssu-humane/HerO,供后续研究使用。

原文摘要 · Abstract (English)

To tackle the AVeriTeC shared task hosted by the FEVER-24, we introduce a system that only employs publicly available large language models (LLMs) for each step of automated fact-checking, dubbed the Herd of Open LLMs for verifying real-world claims (HerO). For evidence retrieval, a language model is used to enhance a query by generating hypothetical fact-checking documents. We prompt pretrained and fine-tuned LLMs for question generation and veracity prediction by crafting prompts with retrieved in-context samples. HerO achieved 2nd place on the leaderboard with the AVeriTeC score of 0.57, suggesting the potential of open LLMs for verifying real-world claims. For future research, we make our code publicly available at https://github.com/ssu-humane/HerO.

事实核查开源模型大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。