arXiv:2410.23850cs.CL2024-10被引 39

评测自动文本事实核查系统,要求判断真假并召回高质量证据。

The Automated Verification of Textual Claims (AVeriTeC) Shared Task

  • 设计多源证据检索与真假判定任务,支持搜索引擎与知识库双路径。
  • 21个参赛方案中18个超越基线,冠军团队达63%的AVeriTeC评分。
  • 适用于事实核查、信息可信度评估等方向的研究者与开发者。

自动化文本事实核查共享任务(AVeriTeC)要求参与者检索真实世界中由事实核查员验证过的陈述的证据,并判断其真伪。证据可通过搜索引擎或组织方提供的知识库获取。提交结果以AVeriTeC得分评估,只有当判断结论正确且召回证据质量达标时才视为准确验证。本次共享任务共收到21份提交,其中18份超过基线表现。冠军团队TUDA_MAI获得63%的AVeriTeC得分。本文介绍该共享任务的设计,呈现全部结果,并总结关键经验与启示。

原文摘要 · Abstract (English)

The Automated Verification of Textual Claims (AVeriTeC) shared task asks participants to retrieve evidence and predict veracity for real-world claims checked by fact-checkers. Evidence can be found either via a search engine, or via a knowledge store provided by the organisers. Submissions are evaluated using AVeriTeC score, which considers a claim to be accurately verified if and only if both the verdict is correct and retrieved evidence is considered to meet a certain quality threshold. The shared task received 21 submissions, 18 of which surpassed our baseline. The winning team was TUDA_MAI with an AVeriTeC score of 63%. In this paper we describe the shared task, present the full results, and highlight key takeaways from the shared task.

事实核查信息验证共享任务证据检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。