用动态词向量更新词典,实时识别新型仇恨语言。
Evolving Hate Speech Online: An Adaptive Framework for Detection and Mitigation
- 基于词向量自动更新词典,适应新出现的侮辱性词汇。
- 融合BERT与词典方法,对主流数据集准确率达95%。
- 适合需要持续防御网络仇恨言论的平台使用。
社交媒体平台的普及加剧了仇恨言论的传播,尤其针对弱势群体。现有自动识别和屏蔽有毒语言的方法依赖预设词典,属于被动响应型,难以应对新出现的歧视性用语。为此,本文提出一种自适应方法,利用词向量动态更新词典,并构建结合BERT与词典策略的混合模型,可有效识别包括故意拼写错误在内的新型毒言。该模型在多数主流数据集上达到95%的准确率。本研究对提升在线环境安全、主动更新检测词库具有重要意义。内容警告:本文包含可能引发不适的仇恨言论实例。
原文摘要 · Abstract (English)
The proliferation of social media platforms has led to an increase in the spread of hate speech, particularly targeting vulnerable communities. Unfortunately, existing methods for automatically identifying and blocking toxic language rely on pre-constructed lexicons, making them reactive rather than adaptive. As such, these approaches become less effective over time, especially when new communities are targeted with slurs not included in the original datasets. To address this issue, we present an adaptive approach that uses word embeddings to update lexicons and develop a hybrid model that adjusts to emerging slurs and new linguistic patterns. This approach can effectively detect toxic language, including intentional spelling mistakes employed by aggressors to avoid detection. Our hybrid model, which combines BERT with lexicon-based techniques, achieves an accuracy of 95% for most state-of-the-art datasets. Our work has significant implications for creating safer online environments by improving the detection of toxic content and proactively updating the lexicon. Content Warning: This paper contains examples of hate speech that may be triggering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。