首次评估大模型在多语言仇恨言论检测中的表现
Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection
- 在英、西、德三语环境下测试GPT-3.5等三模型表现
- 非英语场景下翻译增强数据可提升检测准确率
- 揭示模型与数据集固有偏见对敏感话题误判的影响
识别网络仇恨言论对维护社交媒体安全至关重要。尽管大语言模型(LLMs)在社交分析中展现出潜力,但在多语言仇恨言论检测任务中仍缺乏系统评估。本文首次在英语、西班牙语和德语三种语言中,对GPT-3.5、Flan-T5和Mistral三类模型进行单语与多语设置下的评估,并研究不同提示语种及翻译增强数据在非英语场景中的影响。此外,还分析了模型与数据集固有偏见对敏感话题误判的关联性。
原文摘要 · Abstract (English)
Identifying offensive language is essential for maintaining safety and sustainability in the social media era. Though large language models (LLMs) have demonstrated encouraging potential in social media analytics, they lack thorough evaluation when in offensive language detection, particularly in multilingual environments. We for the first time evaluate multilingual offensive language detection of LLMs in three languages: English, Spanish, and German with three LLMs, GPT-3.5, Flan-T5, and Mistral, in both monolingual and multilingual settings. We further examine the impact of different prompt languages and augmented translation data for the task in non-English contexts. Furthermore, we discuss the impact of the inherent bias in LLMs and the datasets in the mispredictions related to sensitive topics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。