arXiv:2501.16750cs.CRcs.LG2025-01被引 27

评测大模型生成仇恨言论的检测效果,发现新模型越难识别,且可被攻击绕过。

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns

  • 构建7838条大模型生成的仇恨言论数据集,覆盖34类身份群体。
  • 检测器对新大模型生成内容识别率下降,对抗攻击成功率高达96.6%。
  • 揭示大模型驱动仇恨攻击新威胁,适合安全与内容审核研究者关注。

大型语言模型(LLMs)在生成仇恨言论方面的滥用引发广泛关注。尽管仇恨言论检测器在此问题中扮演关键角色,但其对大模型生成内容的检测效果仍不明确。本文提出HateBench框架,用于评估仇恨言论检测器在大模型生成内容上的表现。我们首先构建了一个包含7,838个样本的数据集,由六种主流大模型生成,覆盖34类身份群体,并由三位标注者进行细致标注。随后评估了八种代表性检测器在该数据集上的表现。结果表明,尽管检测器总体上能有效识别大模型生成的仇恨言论,但其性能随大模型版本更新而下降。此外,我们揭示了大模型驱动的仇恨攻击新威胁:攻击者可利用对抗攻击和模型窃取攻击,主动规避检测并自动化网络仇恨活动。最强对抗攻击的成功率达0.966,结合模型窃取攻击后效率提升13-21倍,同时保持良好攻击性能。本研究呼吁学术界与平台管理者加强对此类新兴威胁的防御能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have raised increasing concerns about their misuse in generating hate speech. Among all the efforts to address this issue, hate speech detectors play a crucial role. However, the effectiveness of different detectors against LLM-generated hate speech remains largely unknown. In this paper, we propose HateBench, a framework for benchmarking hate speech detectors on LLM-generated hate speech. We first construct a hate speech dataset of 7,838 samples generated by six widely-used LLMs covering 34 identity groups, with meticulous annotations by three labelers. We then assess the effectiveness of eight representative hate speech detectors on the LLM-generated dataset. Our results show that while detectors are generally effective in identifying LLM-generated hate speech, their performance degrades with newer versions of LLMs. We also reveal the potential of LLM-driven hate campaigns, a new threat that LLMs bring to the field of hate speech detection. By leveraging advanced techniques like adversarial attacks and model stealing attacks, the adversary can intentionally evade the detector and automate hate campaigns online. The most potent adversarial attack achieves an attack success rate of 0.966, and its attack efficiency can be further improved by $13-21\times$ through model stealing attacks with acceptable attack performance. We hope our study can serve as a call to action for the research community and platform moderators to fortify defenses against these emerging threats.

仇恨言论大模型安全检测评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。