研究说话人族裔如何影响大模型识别仇恨言论,发现方言特征比明示身份更易导致误判。
Who Speaks Matters: Analysing the Influence of the Speaker's Ethnicity on Hate Classification
- 通过注入显性身份语句和隐性方言特征,测试模型对说话人背景的敏感度。
- 隐性方言特征导致模型判断翻转比例高于显性身份提示,且不同族裔差异明显。
- 更大模型表现更稳健,提示部署高风险任务需谨慎评估偏见风险。
大型语言模型(LLMs)在内容审核中具有规模化检测仇恨言论的潜力,但其对边缘群体和方言存在脆弱性和偏见。本文研究在输入中引入说话人族裔的显性与隐性标记时,模型在仇恨言论分类上的鲁棒性。显性标记为直接提及语言身份的短语,隐性标记则为方言特征。分析显示,三种LLM和一种LM在五种语言身份下表现不一:隐性方言特征引发的判断翻转频率高于显性身份标记;不同族裔的翻转比例存在差异;且模型规模越大,鲁棒性越强。结果表明,将此类模型用于仇恨言论检测等高风险任务时需格外谨慎。
原文摘要 · Abstract (English)
Large Language Models (LLMs) offer a lucrative promise for scalable content moderation, including hate speech detection. However, they are also known to be brittle and biased against marginalised communities and dialects. This requires their applications to high-stakes tasks like hate speech detection to be critically scrutinized. In this work, we investigate the robustness of hate speech classification using LLMs particularly when explicit and implicit markers of the speaker's ethnicity are injected into the input. For explicit markers, we inject a phrase that mentions the speaker's linguistic identity. For the implicit markers, we inject dialectal features. By analysing how frequently model outputs flip in the presence of these markers, we reveal varying degrees of brittleness across 3 LLMs and 1 LM and 5 linguistic identities. We find that the presence of implicit dialect markers in inputs causes model outputs to flip more than the presence of explicit markers. Further, the percentage of flips varies across ethnicities. Finally, we find that larger models are more robust. Our findings indicate the need for exercising caution in deploying LLMs for high-stakes tasks like hate speech detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。