利用已有数据挖掘隐性仇恨言论,提升检测泛化能力。
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
- 通过关键词分析定位隐藏的隐性仇恨内容
- 用大模型重标注并扩充数据,F1提升12.9点
- 适合研究隐性仇恨检测与数据增强的学者
隐性仇恨言论已成为社交媒体平台的重大挑战。尽管以往研究多聚焦于显性有害言论,但对隐蔽、微妙形式的仇恨言论检测需求日益迫切。基于词典分析,我们假设隐性仇恨言论已存在于公开的有害言论数据集中,但未被标注人员识别或标注。此外,众包数据集因任务复杂且受标注者主观影响,常出现误标。本文提出一种方法,利用现有有害言论数据集提升隐性仇恨言论检测的泛化能力。方法包含三个关键步骤:关键样本识别、重标注及基于 Llama-3 70B 和 GPT-4o 的数据增强。实验表明,该方法显著提升了隐性仇恨检测效果,相比基线提升 12.9 点 F1 分数。
原文摘要 · Abstract (English)
Implicit hate speech has recently emerged as a critical challenge for social media platforms. While much of the research has traditionally focused on harmful speech in general, the need for generalizable techniques to detect veiled and subtle forms of hate has become increasingly pressing. Based on lexicon analysis, we hypothesize that implicit hate speech is already present in publicly available harmful speech datasets but may not have been explicitly recognized or labeled by annotators. Additionally, crowdsourced datasets are prone to mislabeling due to the complexity of the task and often influenced by annotators' subjective interpretations. In this paper, we propose an approach to address the detection of implicit hate speech and enhance generalizability across diverse datasets by leveraging existing harmful speech datasets. Our method comprises three key components: influential sample identification, reannotation, and augmentation using Llama-3 70B and GPT-4o. Experimental results demonstrate the effectiveness of our approach in improving implicit hate detection, achieving a +12.9-point F1 score improvement compared to the baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。