用双类提示生成提升印尼性别仇恨言论检测效果
Dual-Class Prompt Generation: Enhancing Indonesian Gender-Based Hate Speech Detection through Data Augmentation
- 提出双类提示生成法,同时利用仇恨与非仇恨样本增强数据
- 达88.5%准确率和88.1% F1分数,优于单类与回译方法
- 适合研究仇恨言论检测与数据增强的中文读者
印尼社交媒体中的性别仇恨言论检测因标注数据有限而困难。尽管二分类仇恨言论识别已有进展,但更具粒度的性别目标型仇恨言论研究不足,主要受限于类别不平衡。本文比较三种数据增强技术在印尼性别仇恨言论检测中的表现:回译、单类提示生成(仅使用仇恨样本)以及提出的双类提示生成(同时使用仇恨与非仇恨样本)。实验显示,所有增强方法均提升分类性能,其中双类方法表现最佳(随机森林模型下准确率88.5%,F1分数88.1%)。语义相似性分析表明,双类生成内容最富创新性;T-SNE可视化证实这些样本占据独特特征空间区域,同时保持类别特性。结果说明,融合两类样本有助于语言模型生成更多样且具代表性数据,有效缓解特定仇恨言论检测中的数据稀缺问题。
原文摘要 · Abstract (English)
Detecting gender-based hate speech in Indonesian social media remains challenging due to limited labeled datasets. While binary hate speech classification has advanced, a more granular category like gender-targeted hate speech is understudied because of class imbalance issues. This paper addresses this gap by comparing three data augmentation techniques for Indonesian gender-based hate speech detection. We evaluate backtranslation, single-class prompt generation (using only hate speech examples), and our proposed dual-class prompt generation (using both hate speech and non-hate speech examples). Experiments show all augmentation methods improve classification performance, with our dual-class approach achieving the best results (88.5% accuracy, 88.1% F1-score using Random Forest). Semantic similarity analysis reveals dual-class prompt generation produces the most novel content, while T-SNE visualizations confirm these samples occupy distinct feature space regions while maintaining class characteristics. Our findings suggest that incorporating examples from both classes helps language models generate more diverse yet representative samples, effectively addressing limited data challenges in specialized hate speech detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。