用编码类型提升隐性仇恨言论识别准确率
Cracking the Code: Enhancing Implicit Hate Speech Detection through Coding Classification
- 提出六种编码策略,构建隐性仇恨言论新分类体系
- 在中英文数据集上均显著提升检测效果
- 适合需要精准识别隐性攻击的平台内容审核场景
互联网已成为仇恨言论(HS)的高发地,威胁社会和谐与个人福祉。尽管自动检测方法在显性仇恨言论(ex-HS)上表现良好,但在更隐蔽的隐性仇恨言论(im-HS)识别上仍存在困难。本文提出一种新的im-HS检测分类体系,定义六种编码策略(codetypes)。设计两种融合codetypes的方法:一是直接用大语言模型(LLMs)基于生成响应进行分类;二是将codetypes嵌入编码过程,作为LLM的编码器。实验表明,该方法在中英文数据集上均有效提升im-HS检测性能,验证了其跨语言有效性。
原文摘要 · Abstract (English)
The internet has become a hotspot for hate speech (HS), threatening societal harmony and individual well-being. While automatic detection methods perform well in identifying explicit hate speech (ex-HS), they struggle with more subtle forms, such as implicit hate speech (im-HS). We tackle this problem by introducing a new taxonomy for im-HS detection, defining six encoding strategies named codetypes. We present two methods for integrating codetypes into im-HS detection: 1) prompting large language models (LLMs) directly to classify sentences based on generated responses, and 2) using LLMs as encoders with codetypes embedded during the encoding process. Experiments show that the use of codetypes improves im-HS detection in both Chinese and English datasets, validating the effectiveness of our approach across different languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。