攻击者用反例让AI误判仇恨内容为正常,或把正常内容标记为仇恨。
Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

- 用反例和对抗性理由干扰AI判断,模拟人工审核员的反驳行为。
- 多轮交互下模型误判率显著上升,不同模型对两种攻击方向的敏感度不同。
- 现有防御方法可缓解但无法根除风险,需针对性设计防护机制。
大型语言模型(LLMs)越来越多地用于仇恨言论审核,通常在人机协作流程中由审核员提供反馈后作出最终决策。此类反馈引入了两种操纵方向:将仇恨内容淡化为正常内容(whitewashing),或将正常内容污名化为仇恨内容(smearing)。本研究考察了初始判断正确的模型在面对标注风格反驳时的脆弱性,并分析两种操纵方向的攻击效果差异。我们提出一种重审协议,通过决策边界扰动和对抗性推理理由扩展直接矛盾。在两个仇恨言论数据集上对多个LLM进行实验,结果表明标注风格反驳会显著降低审核性能,且在多轮场景下影响更明显。研究还发现,不同模型在两种攻击方向间存在稳定的、特有的不对称性,揭示出不同的方向性脆弱模式。显式推理提示和防御指令虽能减轻影响,但无法彻底消除。这些发现强调了在人机协同审核流程中需要具备方向感知的防护措施和专门的反馈鲁棒性评估。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used for hate speech moderation, often within human--AI workflows in which reviewers provide feedback before a final decision. Such feedback introduces two manipulation directions: whitewashing hateful content as normal and smearing normal content as hateful. This study examines the susceptibility of initially correct model judgments to annotator-style rebuttals and analyzes whether attack effectiveness differs across manipulation directions. We introduce a rejudge protocol that extends direct contradiction with decision-boundary perturbations and adversarial rationales. Experiments with multiple LLMs on two hate speech datasets show that annotator-style rebuttals substantially degrade moderation performance, with stronger effects in multi-turn settings. The results further reveal stable, model-specific asymmetries between whitewashing and smearing across attack configurations, indicating distinct directional vulnerability patterns. Explicit reasoning prompts and defensive instructions reduce these effects but do not eliminate them. These findings highlight the need for direction-aware safeguards and dedicated feedback-robustness evaluation in human--AI moderation workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。