构建多语言极端内容数据集,揭示标注偏差对模型的影响
Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection
- 构建英法阿三语极端内容数据集,标注激进化程度与行动号召
- 发现标注者间存在显著分歧,影响模型性能与公平性
- 用合成数据探究社会属性如何影响标注与预测结果
在线平台上的极端内容传播带来暴力煽动与极端思想扩散等重大风险。尽管研究持续进行,现有数据集和模型仍难以应对多语言、多样化数据的复杂性。为此,我们发布一个公开可用的多语言数据集,涵盖英语、法语和阿拉伯语,标注了激进化水平、行动号召及命名实体。数据集经过去标识化处理,在保护隐私的同时保留上下文信息。除提供数据集外,我们分析了标注过程中的偏见与标注者分歧,探讨其对模型性能的影响。此外,利用合成数据研究社会人口特征对标注模式和模型预测的影响。本工作全面审视了构建鲁棒极端内容检测数据集的挑战与机遇,强调模型开发中公平性与透明性的关键作用。
原文摘要 · Abstract (English)
The proliferation of radical content on online platforms poses significant risks, including inciting violence and spreading extremist ideologies. Despite ongoing research, existing datasets and models often fail to address the complexities of multilingual and diverse data. To bridge this gap, we introduce a publicly available multilingual dataset annotated with radicalization levels, calls for action, and named entities in English, French, and Arabic. This dataset is pseudonymized to protect individual privacy while preserving contextual information. Beyond presenting our freely available dataset, we analyze the annotation process, highlighting biases and disagreements among annotators and their implications for model performance. Additionally, we use synthetic data to investigate the influence of socio-demographic traits on annotation patterns and model predictions. Our work offers a comprehensive examination of the challenges and opportunities in building robust datasets for radical content detection, emphasizing the importance of fairness and transparency in model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。