对比三种方法检测波斯语推特不文明内容,发现模型效果不如人工标注。
Old wine in old glasses: Comparing computational and qualitative methods in identifying incivility on Persian Twitter during the #MahsaAmini movement
- 用人类标注、ParsBERT和ChatGPT三种方式分析伊朗#MahsaAmini运动推文
- ParsBERT在识别仇恨言论上显著优于七个测试的ChatGPT模型
- 英文提示对ChatGPT输出无明显影响,且难以处理隐晦与明确不文明内容
本文比较了三种识别波斯语推文不文明内容的方法:人工定性编码、基于ParsBERT的有监督学习,以及大语言模型(ChatGPT)。基于47,278条来自伊朗#MahsaAmini运动的推文,评估了各方法的准确性和效率。结果表明,ParsBERT在识别仇恨言论方面显著优于七个测试的ChatGPT模型。同时发现,ChatGPT不仅难以处理语义微妙的不文明内容,对明显不文明文本也表现不佳;使用英文或波斯语提示对其输出无显著影响。研究详细对比了三类方法,明确了其在低资源语言环境下的优劣,为仇恨言论分析提供参考。
原文摘要 · Abstract (English)
This paper compares three approaches to detecting incivility in Persian tweets: human qualitative coding, supervised learning with ParsBERT, and large language models (ChatGPT). Using 47,278 tweets from the #MahsaAmini movement in Iran, we evaluate the accuracy and efficiency of each method. ParsBERT substantially outperforms seven evaluated ChatGPT models in identifying hate speech. We also find that ChatGPT struggles not only with subtle cases but also with explicitly uncivil content, and that prompt language (English vs. Persian) does not meaningfully affect its outputs. The study provides a detailed comparison of these approaches and clarifies their strengths and limitations for analyzing hate speech in a low-resource language context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。