arXiv:2411.19017cs.CLcs.LG2024-11综述被引 10

系统梳理低资源语言仇恨言论检测的研究现状与挑战

A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages

  • 综述全球低资源语言仇恨言论检测的现有数据集、特征与技术
  • 指出当前研究因缺乏数据而进展缓慢,尤其在非英语语境下
  • 适合关注多语言AI安全、社会媒体治理的研究者参考

过去十年社交媒体的广泛影响改变了人们的交流方式。社交媒体提供的匿名性与互联网的易访问性助长了仇恨言论的传播。仇恨言论的用语随时代演变,给政策制定者和研究者识别工作带来困难。随着越来越多用户使用母语交流,低资源语言中的仇恨言论也日益增多。尽管英语相关方法已有一定认知,但因缺少数据集与公开数据,针对这些语言的研究仍严重不足。本文详细综述了全球范围内低资源语言仇恨言论检测的现状,涵盖可用数据集、特征与技术,并讨论现有综述、相关概念重叠、研究挑战与机遇。

原文摘要 · Abstract (English)

The expanding influence of social media platforms over the past decade has impacted the way people communicate. The level of obscurity provided by social media and easy accessibility of the internet has facilitated the spread of hate speech. The terms and expressions related to hate speech gets updated with changing times which poses an obstacle to policy-makers and researchers in case of hate speech identification. With growing number of individuals using their native languages to communicate with each other, hate speech in these low-resource languages are also growing. Although, there is awareness about the English-related approaches, much attention have not been provided to these low-resource languages due to lack of datasets and online available data. This article provides a detailed survey of hate speech detection in low-resource languages around the world with details of available datasets, features utilized and techniques used. This survey further discusses the prevailing surveys, overlapping concepts related to hate speech, research challenges and opportunities.

仇恨言论低资源语言综述社会媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。