跨语言LGBTQIA+仇恨言论检测存在显著差异,翻译会丢失关键语义。
"Is Hate Lost in Translation?": Evaluation of Multilingual LGBTQIA+ Hate Speech Detection
- 对比原生文本与机器翻译后的文本,评估大模型检测能力
- 英语表现最佳,英-泰混用语境下效果最差,翻译提升不一
- 揭示翻译难以捕捉文化语境,适合多语言内容安全研究者
本文研究大语言模型在多种语言(包括英语、意大利语、中文及英-泰混用)中检测LGBTQIA+仇恨言论的挑战,考察机器翻译对仇恨言论语义的影响。通过零样本和微调的GPT模型进行实验,发现:(1) 英语表现最优,英-泰混用场景最差;(2) 微调在各语言中均提升性能,而翻译结果则表现不一。通过对原始文本与机器翻译文本的对比实验及定性错误分析,揭示了语言中社会文化语境的复杂性可能无法被自动翻译完整保留。
原文摘要 · Abstract (English)
This paper explores the challenges of detecting LGBTQIA+ hate speech of large language models across multiple languages, including English, Italian, Chinese and (code-switched) English-Tamil, examining the impact of machine translation and whether the nuances of hate speech are preserved across translation. We examine the hate speech detection ability of zero-shot and fine-tuned GPT. Our findings indicate that: (1) English has the highest performance and the code-switching scenario of English-Tamil being the lowest, (2) fine-tuning improves performance consistently across languages whilst translation yields mixed results. Through simple experimentation with original text and machine-translated text for hate speech detection along with a qualitative error analysis, this paper sheds light on the socio-cultural nuances and complexities of languages that may not be captured by automatic translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。