针对不断演化的毒害内容伪装,提出持续学习检测框架,提升识别鲁棒性。
ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations
- 用大模型动态补充被篡改文本的语义与毒性线索,增强理解。
- 通过强化关键特征、抑制无关特征,构建更稳健的判别边界。
- 适合需要长期抵御新型毒害内容的平台安全系统使用。
毒性检测有助于遏制在线社交中仇恨评论、帖子和消息等有毒内容的传播,维护健康网络环境。然而,恶意用户持续开发逃避性扰动以伪装有毒内容并绕过检测器。传统检测方法静态固定,难以应对不断演进的规避策略。因此,持续学习成为动态更新检测能力的合理方案。但扰动间差异阻碍了对扰动文本的持续学习;更重要的是,扰动引入的噪声会扭曲语义,干扰理解,同时损害关键特征学习,使检测对扰动过于敏感。这些因素加剧了对抗演化扰动的持续学习挑战。本文提出ContiGuard,首个面向时间演化扰动文本的持续毒性检测框架,使检测器能持续更新能力并保持对抗演化扰动的持久韧性。具体而言,为提升理解,提出基于大模型的语义增强策略,动态融入大模型挖掘出的可能语义和毒性线索至扰动文本中。为减少非关键特征影响、突出关键特征,设计可区分性驱动的特征学习策略,强化判别性特征,抑制非判别性特征,塑造稳健的分类边界。
原文摘要 · Abstract (English)
Toxicity detection mitigates the dissemination of toxic content (e.g., hateful comments, posts, and messages within online social actions) to safeguard a healthy online social environment. However, malicious users persistently develop evasive perturbations to disguise toxic content and evade detectors. Traditional detectors or methods are static over time and are inadequate in addressing these evolving evasion tactics. Thus, continual learning emerges as a logical approach to dynamically update detection ability against evolving perturbations. Nevertheless, disparities across perturbations hinder the detector's continual learning on perturbed text. More importantly, perturbation-induced noises distort semantics to degrade comprehension and also impair critical feature learning to render detection sensitive to perturbations. These amplify the challenge of continual learning against evolving perturbations. In this work, we present ContiGuard, the first framework tailored for continual learning of the detector on time-evolving perturbed text (termed continual toxicity detection) to enable the detector to continually update capability and maintain sustained resilience against evolving perturbations. Specifically, to boost the comprehension, we present an LLM-powered semantic enriching strategy, where we dynamically incorporate possible meaning and toxicity-related clues excavated by LLM into the perturbed text to improve the comprehension. To mitigate non-critical features and amplify critical ones, we propose a discriminability-driven feature learning strategy, where we strengthen discriminative features while suppressing the less-discriminative ones to shape a robust classification boundary for detection...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。