arXiv:2602.09343cs.AIcs.CL2026-02被引 1

用形式化推理增强检测系统,抵御恶意否定攻击。

Not-in-Perspective: Towards Shielding Google's Perspective API Against Adversarial Negation Attacks

  • 设计推理包装层,前置过滤与后置校正双重机制
  • 在对抗数据集上使毒性识别准确率显著提升
  • 适合安全敏感场景的在线内容审核系统使用

社交媒体中仇恨言论泛滥催生了对自动化毒性检测的需求。现有基于机器学习的检测系统易受逻辑类攻击影响,如语句中的否定表达。本文提出一套基于形式化推理的方法,作为现有机器学习模型的封装层,在预处理和后处理阶段发挥作用。该方法能有效缓解否定攻击带来的误判问题,显著提升毒性评分的准确性与鲁棒性。我们在多个机器学习模型上评估了不同变体的包装方案,并在包含否定攻击的对抗数据集上进行测试。实验结果表明,混合(形式推理+机器学习)方法在面对纯统计模型时表现更优。

原文摘要 · Abstract (English)

The rise of cyberbullying in social media platforms involving toxic comments has escalated the need for effective ways to monitor and moderate online interactions. Existing solutions of automated toxicity detection systems, are based on a machine or deep learning algorithms. However, statistics-based solutions are generally prone to adversarial attacks that contain logic based modifications such as negation in phrases and sentences. In that regard, we present a set of formal reasoning-based methodologies that wrap around existing machine learning toxicity detection systems. Acting as both pre-processing and post-processing steps, our formal reasoning wrapper helps alleviating the negation attack problems and significantly improves the accuracy and efficacy of toxicity scoring. We evaluate different variations of our wrapper on multiple machine learning models against a negation adversarial dataset. Experimental results highlight the improvement of hybrid (formal reasoning and machine-learning) methods against various purely statistical solutions.

内容审核对抗攻击推理系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。