arXiv:2503.22280cs.CL2025-03EMNLP被引 9

构建多语言事实核查声明聚类数据集,提升跨语言查证效率

MultiClaimNet: A Massively Multilingual Dataset of Fact-Checked Claim Clusters

  • 基于相似性匹配自动聚类86种语言的声明,人工干预极少
  • 最大数据集含85.3万条78种语言的已验证声明
  • 适合研究跨语言信息去重与自动化事实核查的学者

在事实核查领域,同一事实常以不同表述在多平台、多语言中重复出现,造成冗余。为减少重复,现有方法尝试检索已有核查结果,但未验证声明数量激增、核查数据库膨胀,亟需更高效方案。本文提出一种将讨论相同事实的声明聚类的方法,以提升检索与验证效率。由于缺乏合适数据集,该研究受限。为此,我们构建了MultiClaimNet——包含三个多语言声明聚类数据集,涵盖86种语言、多种主题。小规模数据集基于已有匹配数据集生成;大规模数据集通过近似最近邻检索生成候选对,并利用大语言模型自动标注相似性,最终形成含85.3万条声明、覆盖78种语言的数据集。我们还测试了多种聚类技术与句子嵌入模型,建立基准性能。本工作为可扩展的声明聚类提供了坚实基础,助力高效事实核查流程。

原文摘要 · Abstract (English)

In the context of fact-checking, claims are often repeated across various platforms and in different languages, which can benefit from a process that reduces this redundancy. While retrieving previously fact-checked claims has been investigated as a solution, the growing number of unverified claims and expanding size of fact-checked databases calls for alternative, more efficient solutions. A promising solution is to group claims that discuss the same underlying facts into clusters to improve claim retrieval and validation. However, research on claim clustering is hindered by the lack of suitable datasets. To bridge this gap, we introduce \textit{MultiClaimNet}, a collection of three multilingual claim cluster datasets containing claims in 86 languages across diverse topics. Claim clusters are formed automatically from claim-matching pairs with limited manual intervention. We leverage two existing claim-matching datasets to form the smaller datasets within \textit{MultiClaimNet}. To build the larger dataset, we propose and validate an approach involving retrieval of approximate nearest neighbors to form candidate claim pairs and an automated annotation of claim similarity using large language models. This larger dataset contains 85.3K fact-checked claims written in 78 languages. We further conduct extensive experiments using various clustering techniques and sentence embedding models to establish baseline performance. Our datasets and findings provide a strong foundation for scalable claim clustering, contributing to efficient fact-checking pipelines.

事实核查多语言聚类数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。