arXiv:2503.15220cs.CL2025-03Conference of the …被引 6

提升跨语言事实核查中对多语种声明的识别能力

Entity-aware Cross-lingual Claim Detection for Automated Fact-checking

  • 利用实体信息增强多语言模型的跨语言泛化能力
  • 在27种语言上均实现性能提升,未见语言也表现稳健
  • 适合需要处理多语种社交媒体信息的核查系统

在自动化事实核查中,识别需要验证的声明是一项关键任务,尤其面对社交媒体上泛滥的虚假信息。尽管已有显著进展,但在处理在线话语中常见的多语言数据方面仍面临挑战。近期工作聚焦于微调预训练多语言模型以应对这一问题。虽然这些模型能处理多种语言,但其在社交媒体传播声明上的跨语言知识迁移能力仍待深入探索。本文提出EX-Claim,一种基于实体感知的跨语言声明检测模型,能有效处理多语言声明。该模型利用命名实体识别与实体链接技术提取的实体信息,在训练过程中提升已见及未见语言的语言层面性能。在三个来自不同社交媒体平台的数据集上进行的广泛实验表明,所提模型在27种语言中均表现出色,并在训练中见过与未见过的语言间展现出强健的知识迁移能力。

原文摘要 · Abstract (English)

Identifying claims requiring verification is a critical task in automated fact-checking, especially given the proliferation of misinformation on social media platforms. Despite notable progress, challenges remain-particularly in handling multilingual data prevalent in online discourse. Recent efforts have focused on fine-tuning pre-trained multilingual language models to address this. While these models can handle multiple languages, their ability to effectively transfer cross-lingual knowledge for detecting claims spreading on social media remains under-explored. In this paper, we introduce EX-Claim, an entity-aware cross-lingual claim detection model that generalizes well to handle multilingual claims. The model leverages entity information derived from named entity recognition and entity linking techniques to improve the language-level performance of both seen and unseen languages during training. Extensive experiments conducted on three datasets from different social media platforms demonstrate that our proposed model stands out as an effective solution, demonstrating consistent performance gains across 27 languages and robust knowledge transfer between languages seen and unseen during training.

跨语言事实核查实体感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。