arXiv:2605.15886cs.CL2026-05

构建俄政府多模态政治演讲数据集,支持跨语言时空分析。

Linked Multi-Modal Data on Russian Domestic and Foreign Policy Speeches

论文配图:Linked Multi-Modal Data on Russian Domestic and Foreign Policy Speeches
图 1 · 摘自论文原文
  • 链接俄克里姆林宫与外交部数十年官方演讲的文本、图像与元数据。
  • 提供中俄双语文本、图像、标注及唯一标识符,实现跨模态对齐。
  • 适合研究威权政治传播或大模型在政论领域应用的研究者。

本文提出一个关于俄罗斯政府多模态政治传播的互连数据集,弥补威权政治语境下社交文本与图像数据稀缺的问题。数据集包含克里姆林宫和俄罗斯外交部高级官员多年来的官方演讲,每篇演讲均提供俄语与英语文本、相关图像及字幕(如有),以及统一的元数据,包括日期、发言人、(地理)位置和官方内容标签。通过唯一标识符将图像与演讲关联,并对中俄文文本版本进行对齐。我们进一步利用基于Transformer的多模态主题建模生成主题标注,并由俄罗斯政治专家验证。该数据资源支持多模态、多语言、时间与空间维度的政治传播分析,为社会科学与大语言模型在政治领域的应用提供重要测试平台。

原文摘要 · Abstract (English)

This paper introduces a dataset of interlinked multimodal political communications from the Russian government, addressing persistent deficiencies in the availability of social text- and image-based data for authoritarian politics contexts. The dataset comprises two large corpora of official speeches delivered by senior actors within the Kremlin and the Russian Ministry of Foreign Affairs over multiple decades. For each speech, we provide Russian- and English-language texts, associated images and captions where available, and harmonized metadata including (e.g.) dates, speakers, (geo)locations, and official government content tags. Unique identifiers link images to speeches and align Russian and English versions of the same communication texts. We further augment these linked datasets with validated topical annotations for both speech texts and speech images, which are generated via transformer-based multimodal topic modeling and refined by a Russian politics expert. The resulting data resources support multimodal, multilingual, temporal, and/or spatial analyses of (authoritarian) political communication and offer a valuable testbed for social science research and large language model (LLM) applications in political domains.

多模态政治传播数据集大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。