arXiv:2509.12672cs.CL2025-09被引 1

找出聊天机器人生成内容下毒性分类器的脆弱点并修复,提升公平性与抗攻击能力

Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content

  • 通过可解释性技术定位分类器中易被攻击的神经元组件
  • 抑制脆弱神经元后,模型在对抗攻击下的准确率显著提升
  • 发现不同人群的脆弱性差异,助力更公平的有害内容检测

由于大型语言模型(LLM)的广泛应用,在线机器生成内容激增,给内容审核系统带来新挑战。传统审核分类器通常基于人类文本训练,面对由LLM生成的内容或对抗攻击时容易误判。现有防御策略多为被动响应,依赖对抗训练或外部检测模型。本文针对微调后的BERT和RoBERTa分类器,利用对抗攻击技术识别其脆弱组件,提出基于机制可解释性的新型修复策略。研究覆盖多个涉及少数群体的数据集,发现模型存在对性能关键或易受攻击的特定神经头,抑制后者可有效提升对抗攻击下的表现。同时揭示不同人口群体间的脆弱性差异,暴露了训练中的公平性与鲁棒性缺口,为构建更具包容性的毒性检测模型提供依据。

原文摘要 · Abstract (English)

The volume of machine-generated content online has grown dramatically due to the widespread use of Large Language Models (LLMs), leading to new challenges for content moderation systems. Conventional content moderation classifiers, which are usually trained on text produced by humans, suffer from misclassifications due to LLM-generated text deviating from their training data and adversarial attacks that aim to avoid detection. Present-day defence tactics are reactive rather than proactive, since they rely on adversarial training or external detection models to identify attacks. In this work, we aim to identify the vulnerable components of toxicity classifiers that contribute to misclassification, proposing a novel strategy based on mechanistic interpretability techniques. Our study focuses on fine-tuned BERT and RoBERTa classifiers, testing on diverse datasets spanning a variety of minority groups. We use adversarial attacking techniques to identify vulnerable circuits. Finally, we suppress these vulnerable circuits, improving performance against adversarial attacks. We also provide demographic-level insights into these vulnerable circuits, exposing fairness and robustness gaps in model training. We find that models have distinct heads that are either crucial for performance or vulnerable to attack and suppressing the vulnerable heads improves performance on adversarial input. We also find that different heads are responsible for vulnerability across different demographic groups, which can inform more inclusive development of toxicity detection models.

毒性检测对抗攻击可解释性公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。