arXiv:2601.03860cs.CL2026-01Conference of the …

构建首个多语种欧洲移民叙事偏见数据集,助力识别极端政治言论。

PartisanLens: A Multilingual Dataset of Hyperpartisan and Conspiratorial Immigration Narratives in European Media

  • 构建1617条西语/意语/葡语新闻标题的多维度标注数据集
  • 验证大模型在偏激与阴谋论内容识别上的表现及局限
  • 探索模型模拟不同立场注释者的能力,适合舆情与社会研究者

识别极端偏执叙事和人口替代阴谋论(PRCT)对遏制虚假信息传播至关重要。这类复杂叙事威胁社会凝聚力与公共安全,因极端偏执加剧政治极化,而PRCT可直接煽动现实中的极端暴力。然而现有资源稀缺,多为英文且常孤立分析偏倚、立场与修辞倾向。为此,我们提出 extsc{PartisanLens},首个包含1617条西班牙语、意大利语和葡萄牙语新闻标题的多语言数据集,涵盖多种政治话语维度。我们评估主流大语言模型(LLMs)在该数据集上的分类表现,建立基准。同时检验其作为自动标注工具的可行性,分析其接近人工标注的能力。结果表明其潜力与当前局限。进一步,我们探索通过引入社会经济与意识形态背景,让模型模拟不同注释者视角的注释模式。本研究提供完整资源与评估框架,支持未来在欧洲语境下检测偏执与阴谋论叙事的研究。

原文摘要 · Abstract (English)

Detecting hyperpartisan narratives and Population Replacement Conspiracy Theories (PRCT) is essential to addressing the spread of misinformation. These complex narratives pose a significant threat, as hyperpartisanship drives political polarisation and institutional distrust, while PRCTs directly motivate real-world extremist violence, making their identification critical for social cohesion and public safety. However, existing resources are scarce, predominantly English-centric, and often analyse hyperpartisanship, stance, and rhetorical bias in isolation rather than as interrelated aspects of political discourse. To bridge this gap, we introduce \textsc{PartisanLens}, the first multilingual dataset of \num{1617} hyperpartisan news headlines in Spanish, Italian, and Portuguese, annotated in multiple political discourse aspects. We first evaluate the classification performance of widely used Large Language Models (LLMs) on this dataset, establishing robust baselines for the classification of hyperpartisan and PRCT narratives. In addition, we assess the viability of using LLMs as automatic annotators for this task, analysing their ability to approximate human annotation. Results highlight both their potential and current limitations. Next, moving beyond standard judgments, we explore whether LLMs can emulate human annotation patterns by conditioning them on socio-economic and ideological profiles that simulate annotator perspectives. At last, we provide our resources and evaluation, \textsc{PartisanLens} supports future research on detecting partisan and conspiratorial narratives in European contexts.

偏执叙事多语种数据阴谋论识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。