发现大模型对特定群体仇恨言论易误拒,提出中英翻译缓解策略
Analyzing Bias in False Refusal Behavior of Large Language Models for Hate Speech Detoxification
- 通过跨语言翻译降低模型误拒率
- 针对国籍、宗教等群体的仇恨言论拒绝率更高
- 适合安全检测与内容审核场景使用
尽管大语言模型(LLMs)被广泛应用于仇恨言论净化,但提示词常触发安全警报导致模型拒绝任务。本研究系统分析了仇恨言论净化中的误拒行为,揭示了语境与语言偏差的影响。在英、多语种数据集上评估九个LLMs,结果表明模型对语义毒性更高的输入及特定目标群体(尤其是国籍、宗教、政治立场)更易拒绝。尽管多语种数据集整体误拒率低于英文,但模型仍存在系统性、语言依赖的偏见。基于此,我们提出简单跨翻译策略:将英文仇恨言论翻译为中文再回译,显著降低误拒率且保留原意,提供一种高效轻量的缓解方法。
原文摘要 · Abstract (English)
While large language models (LLMs) have increasingly been applied to hate speech detoxification, the prompts often trigger safety alerts, causing LLMs to refuse the task. In this study, we systematically investigate false refusal behavior in hate speech detoxification and analyze the contextual and linguistic biases that trigger such refusals. We evaluate nine LLMs on both English and multilingual datasets, our results show that LLMs disproportionately refuse inputs with higher semantic toxicity and those targeting specific groups, particularly nationality, religion, and political ideology. Although multilingual datasets exhibit lower overall false refusal rates than English datasets, models still display systematic, language-dependent biases toward certain targets. Based on these findings, we propose a simple cross-translation strategy, translating English hate speech into Chinese for detoxification and back, which substantially reduces false refusals while preserving the original content, providing an effective and lightweight mitigation approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。