arXiv:2601.16309cs.CLcs.SI2026-01

构建跨国家多语言新闻语料库,分析俄乌战争中不同媒体的叙事差异。

A Longitudinal, Multinational, and Multilingual Corpus of News Coverage of the Russo-Ukrainian War

  • 收集246,000篇来自五国十一媒体的战时新闻,覆盖俄、乌、美、英、中
  • 发现媒体通过选择性归因和话题聚焦制造相互矛盾的现实叙述
  • 适合研究国际传播、信息战与跨语言叙事演化,支持实证分析

我们提出DNIPRO,一个涵盖24.6万篇新闻文章的语料库,记录了2022年2月至2024年8月期间俄乌战争的报道,覆盖俄罗斯、乌克兰、美国、英国、中国五个国家的十一家媒体,包含三种语言。该语料库配备详尽元数据及人工标注的立场、情感与主题框架标签,可系统分析地缘政治叙事分歧。探索性分析显示,媒体通过差异化的归因方式与话题选择构建出互不兼容的现实图景,且未直接反驳对方叙事。DNIPRO为叙事演变、跨语言信息流动与碎片化信息生态中隐性矛盾的计算检测提供实证研究支持。

原文摘要 · Abstract (English)

We present DNIPRO, a corpus of 246K news articles from the Russo-Ukrainian war (Feb 2022 -- Aug 2024) spanning eleven outlets across five nation-states (Russia, Ukraine, U.S., U.K., China) and three languages. The corpus features comprehensive metadata and human-evaluated annotations for stance, sentiment, and topical framing, enabling systematic analysis of competing geopolitical narratives. It is uniquely suited for empirical studies of narrative divergence, media framing, and information warfare. Our exploratory analyses reveal how media outlets construct incompatible realities through divergent attribution and topical selection without direct refutation of opposing narratives. DNIPRO empowers empirical research on narrative evolution, cross-lingual information flow, and computational detection of implicit contradictions in fragmented information ecosystems.

新闻语料叙事分析多语言信息战

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。