arXiv:2608.20984cs.CVcs.CL2026-08中稿 · EMNLP

首个面向英国移民叙事的多模态视频数据集,助力理解社交媒体中的移民议题

MigrationNarrate: A Dataset for Detection of Migration Narratives in YouTube Videos

论文配图:MigrationNarrate: A Dataset for Detection of Migration Narratives in YouTube Videos
图 1 · 摘自论文原文
  • 构建包含1115个视频的多模态数据集,采用12类主叙事与53个标签体系标注
  • 基于预训练模型和大语言模型在该数据集上实现基准性能,验证方法可行性
  • 适用于媒体分析、社会舆情研究及政策制定者,推动移民话题的客观解读

叙述是社会传播的核心框架,其检测对理解公共话语至关重要。以往研究已在多个领域探索叙述识别,但移民叙述仍严重缺乏专门标注数据集,且随着公众传播向视频平台迁移,多模态信号成为主流,相关视频中的叙述研究仍处于空白。为此,我们推出首个针对英国移民叙事的多模态数据集——MigrationNarrate,包含1,115条YouTube视频转录文本,采用两级分类体系:12类移民主叙事与53个具体叙事标签。本文详述数据集的设计、采集与标注过程,并利用预训练编码器与开源/闭源大语言模型进行基准测试。最后通过全面误差分析为未来研究提供方向。

原文摘要 · Abstract (English)

Narratives are central to how social communication is framed, making their detection critical for understanding and analysing public discourse. Prior work has explored narrative detection and extraction across diverse domains; however, migration narratives remain significantly understudied, primarily due to the absence of dedicated annotated datasets. Furthermore, public communication has recently shifted towards video-centric platforms, where narratives are conveyed through multimodal signals and consumed at scale. Despite this shift, narratives in videos remain largely unexplored. To bridge these gaps, we introduce MigrationNarrate, the first multimodal dataset for detection of migration narratives in the UK, consisting of 1,115 YouTube video transcripts annotated using a two-level taxonomy of 12 migration super-narratives and 53 narrative labels. This paper details the dataset design, collection, and annotations; together with benchmark results using a combination of pre-trained encoder models and both open- and closed-source Large Language Models. Finally, a thorough error analysis offers insights for future work.

多模态数据集移民叙事视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。