针对毒性内容演变,提出安全感知的动态检测与精准更新机制。
DriftGuard: Safety-Aware Multi-Monitor Detection and Selective Adaptation for Evolving Toxicity Moderation
- 多监控器检测局部安全风险,而非仅依赖全局分布变化。
- 硬混合更新策略提升召回率至0.8777(Civil Comments)和0.8523(DynaHate)。
- 适合需要持续防御新型毒性行为的平台安全系统使用。
自动化毒性审核系统在动态在线环境中运行,有害行为通过编码语言、目标转移及对抗性适应不断演化。现有漂移检测方法多关注全局分布变化,但可能忽略局部危害子空间或高风险模型错误区域的安全相关漂移。本文提出DriftGuard,一种安全感知的自适应审核框架,结合多监控器漂移检测与选择性模型更新。该框架追踪全局文本漂移、身份伤害漂移、模型不确定性、毒性风险漂移及误报风险漂移。当检测到安全相关变化时,采用硬混合适配集进行模型更新,优先处理潜在误漏检、身份相关高风险样本、误报风险样本及不确定边界案例。在Civil Comments时间漂移和Jigsaw-to-DynaHate跨数据集漂移实验中,安全感知监控器识别出全局漂移遗漏的风险。硬混合适配显著提升毒性召回率与准确率,相较无更新与随机平衡基线,在Civil Comments上召回率达0.8777,在DynaHate上从0.7107升至0.8523。自举分析显示,DynaHate上毒性召回率提升0.1418,误漏检率下降0.0781。整体上,DriftGuard将安全感知漂移检测与针对性轻量更新结合,实现更鲁棒的自适应毒性审核。
原文摘要 · Abstract (English)
Automated toxicity moderation systems operate in dynamic online environments where harmful behavior evolves through coded language, shifting targets, and strategic adaptation to enforcement. Existing drift detection methods often focus on global distributional change, but such signals may miss safety-relevant shifts that emerge in localized harm subspaces or high-risk model-error regions. This paper introduces DriftGuard, a safety-aware adaptive moderation framework that combines multi-monitor drift detection with selective model updating. The framework tracks global text drift, identity-harm drift, model uncertainty, toxic-risk drift, and false-negative-risk drift. When safety-relevant change is detected, the model is updated using a hard-mix adaptation set that prioritizes likely false negatives, identity-related high-risk examples, false-positive-risk examples, and uncertain boundary cases. Experiments on Civil Comments temporal shift and Jigsaw-to-DynaHate cross-dataset shift show that safety-aware monitors detect risks missed by global drift alone. Hard-mix adaptation improves toxic recall and accuracy over no-update and random-balanced baselines, raising toxic recall to 0.8777 on Civil Comments and from 0.7107 to 0.8523 on DynaHate. Bootstrap analysis further shows stable DynaHate safety gains, with toxic recall increasing by 0.1418 and false-negative prevalence decreasing by 0.0781. Overall, DriftGuard links safety-aware drift detection to targeted, lightweight model updating for more robust adaptive toxicity moderation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。