用大模型检测拉普拉塔河西班牙语仇恨言论,发现其对隐晦表达更敏感。
Exploring Large Language Models for Hate Speech Detection in Rioplatense Spanish
- 用思维链提示让GPT-3.5、Mixtral等大模型分类仇恨言论
- 大模型在识别同性恋/跨性别歧视时准确率高于传统BERT模型
- 适合研究小语种仇恨言论检测的研究者参考
仇恨言论检测面临多种语言变体、俚语、侮辱性词汇及文化细微差别。本文研究大型语言模型在拉普拉塔河西班牙语仇恨言论检测中的表现。通过思维链提示,对比了ChatGPT 3.5、Mixtral和Aya与先进BERT分类器的分类效果。结果表明,尽管大模型精度低于微调过的BERT,但在识别难以捕捉的侮辱性词汇和方言表达方面表现良好,尤其在同性恋/跨性别歧视等高度隐晦的仇恨言论上具有更强敏感性。代码与模型已公开以供后续研究。
原文摘要 · Abstract (English)
Hate speech detection deals with many language variants, slang, slurs, expression modalities, and cultural nuances. This outlines the importance of working with specific corpora, when addressing hate speech within the scope of Natural Language Processing, recently revolutionized by the irruption of Large Language Models. This work presents a brief analysis of the performance of large language models in the detection of Hate Speech for Rioplatense Spanish. We performed classification experiments leveraging chain-of-thought reasoning with ChatGPT 3.5, Mixtral, and Aya, comparing their results with those of a state-of-the-art BERT classifier. These experiments outline that, even if large language models show a lower precision compared to the fine-tuned BERT classifier and, in some cases, they find hard-to-get slurs or colloquialisms, they still are sensitive to highly nuanced cases (particularly, homophobic/transphobic hate speech). We make our code and models publicly available for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。