arXiv:2409.17130cs.CL2024-09中稿 · publication in "18…被引 2

分析孟加拉语社交平台对三大群体的毒性言论,提升识别精度。

Assessing the Level of Toxicity Against Distinct Groups in Bangla Social Media Comments: A Comprehensive Investigation

  • 构建多源孟加拉语毒性评论数据集并人工标注
  • Bangla-BERT模型F1得分达0.8903,表现最优
  • 揭示不同群体遭遇的毒性差异,助力内容安全

社交媒体在现代社会中扮演重要角色,是沟通、思想交流与网络建立的渠道。然而,滥用平台发布从冒犯性言论到仇恨言论的毒性评论,已成为严峻问题。本研究聚焦于识别针对跨性别者、土著群体和移民群体的孟加拉语毒性评论,涵盖多个社交平台来源。研究深入探讨毒性语言的识别与分类过程,并考虑毒性程度的差异:高、中、低。方法包括构建数据集、人工标注,以及使用Bangla-BERT、bangla-bert-base、distil-BERT和Bert-base-multilingual-cased等预训练Transformer模型进行分类。采用准确率、召回率、精确率和F1分数等多种评估指标衡量模型性能。实验结果表明,Bangla-BERT优于其他模型,达到F1-score 0.8903。该研究揭示了孟加拉语社交对话中毒性的复杂性,凸显其对不同人口群体的差异化影响。

原文摘要 · Abstract (English)

Social media platforms have a vital role in the modern world, serving as conduits for communication, the exchange of ideas, and the establishment of networks. However, the misuse of these platforms through toxic comments, which can range from offensive remarks to hate speech, is a concerning issue. This study focuses on identifying toxic comments in the Bengali language targeting three specific groups: transgender people, indigenous people, and migrant people, from multiple social media sources. The study delves into the intricate process of identifying and categorizing toxic language while considering the varying degrees of toxicity: high, medium, and low. The methodology involves creating a dataset, manual annotation, and employing pre-trained transformer models like Bangla-BERT, bangla-bert-base, distil-BERT, and Bert-base-multilingual-cased for classification. Diverse assessment metrics such as accuracy, recall, precision, and F1-score are employed to evaluate the model's effectiveness. The experimental findings reveal that Bangla-BERT surpasses alternative models, achieving an F1-score of 0.8903. This research exposes the complexity of toxicity in Bangla social media dialogues, revealing its differing impacts on diverse demographic groups.

毒性检测孟加拉语社会媒体公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。