构建多语言多文化在线极化数据集,推动跨文化极化研究
POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization
- 构建涵盖22种语言、11万+样本的多事件极化数据集
- 小模型在二元极化检测上表现良好,但类型与表现形式预测能力弱
- 适用于跨文化社会计算、数字极化治理等研究
在线极化正威胁民主对话,但现有计算社会科学大多局限于单一语言、特定文化或事件。我们提出POLAR,一个涵盖22种语言、超过11万实例的多语言、多文化、多事件数据集,数据来自多样在线平台和真实世界事件。极化现象从检测、类型到表现形式三个维度进行标注,并根据不同文化背景适配标注平台。我们开展两项实验:(1) 微调六种预训练小型语言模型;(2) 在少样本和零样本设置下评估多种开源与闭源大模型。结果表明,尽管多数模型在二元极化检测中表现良好,但在预测极化类型与表现形式时性能显著下降。这揭示了极化的高度情境依赖性,凸显了自然语言处理与计算社会科学需发展更鲁棒、可适应的方法。所有资源将公开,以支持全球数字极化研究与干预。
原文摘要 · Abstract (English)
Online polarization poses a growing challenge for democratic discourse, yet most computational social science research remains monolingual, culturally narrow, or event-specific. We introduce POLAR, a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events. Polarization is annotated along three axes, namely detection, type, and manifestation, using a variety of annotation platforms adapted to each cultural context. We conduct two main experiments: (1) fine-tuning six pretrained small language models; and (2) evaluating a range of open and closed large language models in few-shot and zero-shot settings. The results show that, while most models perform well in binary polarization detection, they achieve substantially lower performance when predicting polarization types and manifestations. These findings highlight the complex, highly contextual nature of polarization and demonstrate the need for robust, adaptable approaches in NLP and computational social science. All resources will be released to support further research and effective mitigation of digital polarization globally.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。