arXiv:2510.19628cs.CL2025-10

构建多语言新闻相似性数据集,助力跨语言假消息检测。

CrossNews-UA: A Cross-lingual News Semantic Similarity Benchmark for Ukrainian, Polish, Russian, and English

  • 用众包方式构建可扩展的跨语言新闻相似性标注流程
  • 推出包含乌、波、俄、英四语新闻对的全新数据集
  • 提供基于4W准则的详细标注,适合多语言信息验证研究

在社交媒体和虚假信息快速传播的时代,新闻分析仍是一项关键任务。跨语言新闻比对通过利用不同语言的外部信息源,为信息验证提供了有效途径(Chen and Shu, 2024)。然而,现有跨语言新闻分析数据集(Chen et al., 2022a)依赖记者和专家手动标注,难以扩展且不适应新语言。本文提出一种可扩展、可解释的众包标注流程,构建了新数据集CrossNews-UA,以乌克兰语为核心语言,涵盖与之语言和语境相关的波兰语、俄语和英语新闻对。每对新闻均基于4W标准(谁、什么、哪里、何时)进行语义相似性标注,并附详细理由。我们测试了从传统词袋模型、Transformer架构到大语言模型(LLMs)等多种模型,结果揭示了多语言新闻分析的挑战,并提供了模型表现的深入见解。

原文摘要 · Abstract (English)

In the era of social networks and rapid misinformation spread, news analysis remains a critical task. Detecting fake news across multiple languages, particularly beyond English, poses significant challenges. Cross-lingual news comparison offers a promising approach to verify information by leveraging external sources in different languages (Chen and Shu, 2024). However, existing datasets for cross-lingual news analysis (Chen et al., 2022a) were manually curated by journalists and experts, limiting their scalability and adaptability to new languages. In this work, we address this gap by introducing a scalable, explainable crowdsourcing pipeline for cross-lingual news similarity assessment. Using this pipeline, we collected a novel dataset CrossNews-UA of news pairs in Ukrainian as a central language with linguistically and contextually relevant languages-Polish, Russian, and English. Each news pair is annotated for semantic similarity with detailed justifications based on the 4W criteria (Who, What, Where, When). We further tested a range of models, from traditional bag-of-words, Transformer-based architectures to large language models (LLMs). Our results highlight the challenges in multilingual news analysis and offer insights into models performance.

跨语言新闻分析数据集假消息检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。