构建首个多标签孟加拉语仇恨言论数据集,解决低资源语言检测难题
BOISHOMMO: Holistic Approach for Bangla Hate Speech
- 构建多标签孟加拉语仇恨言论数据集,涵盖种族、性别、宗教等维度
- 包含2000+标注样本,揭示非拉丁文字处理复杂性与多属性共现特征
- 为低资源语言仇恨言论研究提供可复用的基准数据,适合自然语言处理研究者
社交媒体上的仇恨言论是数字社会中一个严重问题,全球研究者对此高度关注。尽管已有大量工作致力于识别与预警系统,但低资源语言(如孟加拉语)仍存在显著空白。主要瓶颈在于缺乏全面的数据集。仇恨言论本身具有多维属性,现有数据集往往忽略其多重滥用特征。为此,本文构建了名为BOISHOMMO的多标签孟加拉语仇恨言论数据集,涵盖种族、性别、宗教、政治等多个类别。该数据集包含超过2000个标注样本,提供了对孟加拉语仇恨言论的细致理解,并凸显处理非拉丁文字的复杂性。通过多种算法评估,验证了模型性能差异,展示了多标签建模的重要性。该数据集将推动低资源语言仇恨言论检测与分析研究的发展。
原文摘要 · Abstract (English)
One of the most alarming issues in digital society is hate speech (HS) on social media. The severity is so high that researchers across the globe are captivated by this domain. A notable amount of work has been conducted to address the identification and alarm system. However, a noticeable gap exists, especially for low-resource languages. Comprehensive datasets are the main problem among the constrained resource languages, such as Bangla. Interestingly, hate speech or any particular speech has no single dimensionality. Similarly, the hate component can simultaneously have multiple abusive attributes, which seems to be missed in the existing datasets. Thus, a multi-label Bangla hate speech dataset named BOISHOMMO has been compiled and evaluated in this work. That includes categories of HS across race, gender, religion, politics, and more. With over two thousand annotated examples, BOISHOMMO provides a nuanced understanding of hate speech in Bangla and highlights the complexities of processing non-Latin scripts. Apart from evaluating with multiple algorithmic approaches, it also highlights the complexities of processing Bangla text and assesses model performance. This unique multi-label approach enriches future hate speech detection and analysis studies for low-resource languages by providing a more nuanced, diverse dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。