arXiv:2608.27844cs.CL2026-08中稿 · EMNLP

首个动态对抗评估框架,模拟用户持续改写规避内容审核。

EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion

论文配图:EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion
图 1 · 摘自论文原文
  • 通过语义聚类迭代优化逃逸策略,兼顾成功率与可读性。
  • 在顶级LLM审核系统上,12轮优化后攻击成功率达80.3%。
  • 适合关注内容安全真实效能的研究者与平台方参考。

现有有害内容检测评估多依赖静态基准,难以反映真实平台中用户根据审核反馈持续修改表达的互动对抗生态,导致离线评分与线上效果存在显著差距。据我们所知,本文提出EvoHarmBench,首个面向内容审核系统的动态对抗评估框架。该框架采用迭代优化循环,在语义聚类层面演化逃逸策略,同时优化逃逸成功率与人类可读性。我们系统评估了广泛应用于现实系统中的基于大模型的防御模型,覆盖5个违规类别下的229个语义子集群,数据源自5,002条真实平台收集的对抗样本。实验表明,即使在领先的商用系统中仍存在显著漏洞:经过12轮优化迭代,在可读性约束下,攻击成功率最高达到80.3%。我们将公开完整基准数据、评估框架与代码,推动内容安全研究从静态评估转向动态对抗评估。

原文摘要 · Abstract (English)

Existing evaluations of harmful content detection rely predominantly on static benchmarks, which struggle to reflect the interactive adversarial ecosystem of real-world content platforms where users continuously revise their expressions in response to moderation feedback. This mismatch creates a significant performance gap between offline benchmark scores and online deployment effectiveness. To the best of our knowledge, we present EvoHarmBench, the first dynamic adversarial evaluation framework for content moderation systems. The framework employs an iterative optimization loop that evolves evasion strategies at the semantic-cluster level, while simultaneously optimizing for evasion success and human readability. We systematically evaluate LLM-based defense models which are widely used in real world moderation systems. The evaluation covers 229 semantic sub-clusters across five violation categories, derived from 5,002 real-world adversarial samples collected from content platforms. Our experiments reveal substantial vulnerabilities even in leading commercial systems: after twelve optimization iterations, the attack success rate under readability constraints reaches 80.3% within SOTA LLM moderators. We will release the full benchmark data, evaluation framework, and code to encourage a shift from static benchmarking toward dynamic adversarial evaluation in content safety research.

内容审核对抗攻击大模型安全动态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。