arXiv:2505.10740cs.CLcs.IR2025-05ACL被引 28

跨语言事实核查检索任务,助力多语种假信息识别。

SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval

  • 设计双赛道:同语种与跨语言的声明检索任务
  • 179人参与,52次提交,31支队伍提交系统论文
  • 揭示有效方法,推动多语言事实核查研究

在线虚假信息的快速传播构成全球性挑战,机器学习被视为潜在解决方案,但多语言环境和低资源语言常被忽视。为此,我们于SemEval 2025组织了多语言声明检索共享任务,旨在识别与社交媒体帖子中新出现的声明在语义上匹配的事实核查声明,涵盖不同语言。任务包含两个子赛道:(1) 同语种赛道,社交媒体帖子与声明语言一致;(2) 跨语言赛道,两者可能分属不同语言。共有179名参与者注册,提交52次测试结果,31支队伍提交了系统论文。本文报告了各赛道表现最佳的系统及最常见、最有效的方法。该共享任务及其数据集与参赛系统为多语言声明检索与自动化事实核查研究提供了宝贵洞见,支持未来研究发展。

原文摘要 · Abstract (English)

The rapid spread of online disinformation presents a global challenge, and machine learning has been widely explored as a potential solution. However, multilingual settings and low-resource languages are often neglected in this field. To address this gap, we conducted a shared task on multilingual claim retrieval at SemEval 2025, aimed at identifying fact-checked claims that match newly encountered claims expressed in social media posts across different languages. The task includes two subtracks: (1) a monolingual track, where social posts and claims are in the same language, and (2) a crosslingual track, where social posts and claims might be in different languages. A total of 179 participants registered for the task contributing to 52 test submissions. 23 out of 31 teams have submitted their system papers. In this paper, we report the best-performing systems as well as the most common and the most effective approaches across both subtracks. This shared task, along with its dataset and participating systems, provides valuable insights into multilingual claim retrieval and automated fact-checking, supporting future research in this field.

事实核查多语言跨语言共享任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。