arXiv:2607.28187cs.AIcs.CR2026-07

简单图像变换可绕过主流AI内容审核,暴露其安全漏洞

Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation

论文配图:Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation
图 1 · 摘自论文原文
  • 用无需梯度的简单变换(如变色、灰度化)攻击审核系统
  • 多种变换下,90%以上测试案例可实现从违规到合规的误判
  • 适合研究安全漏洞或设计多层审核系统的开发者参考

尽管自动化内容审核系统已成为大规模筛查有害内容的关键,但传统任务特定分类器在政策覆盖和上下文理解方面仍有限。近期,基于大模型的商业多模态审核接口被推出,宣称具备更广更强的安全过滤能力。本文对三家主流商业图像审核服务进行了大规模黑盒评估,对比其鲁棒性。通过在多个提供商、数据集、危害类别、感知相似性约束和变换强度下测试七种简单、模型无关的图像变换,发现:(1) 所有三款服务均可通过无需梯度、无需替代模型或目标系统知识的低成本变换成功绕过;(2) 即使是固定变换如颜色反转和灰度化,也能引发人类仍可识别的内容从违规到合规的判定转变;(3) 其鲁棒性在不同数据集和危害类别间差异显著,多模态内容与自残类内容尤为脆弱。结论表明,仅用大模型基础模型接口替换传统分类器,并不能自动提供可靠的安全部署边界。此类系统必须在真实变换条件下评估,并作为分层审核流程中的一环部署,而非独立安全屏障。

原文摘要 · Abstract (English)

While automated content-moderation systems have become essential for screening harmful content at scale, conventional task-specific classifiers often provide limited policy cov- erage and contextual understanding. Recently, commercial multimodal moderation APIs built on large foundation models have been introduced with the promise of providing broader and more capable safety filters. In this work, we analyze whether this shift also yields more robust image moderation. We conduct a large-scale black-box evaluation on three established commercial image-moderation services and compare their robustness. By evaluating seven simple, model-agnostic image transformations across multiple providers, datasets, harm categories, perceptual-similarity constraints, and transformation intensities, we find that: (1) all three commercial services can be bypassed using inexpensive image transformations that require no gradients, surrogate models, or knowledge of the target system; (2) even fixed transformations such as color inversion and grayscale conversion induce unsafe-to-safe decision changes while preserving content that remains recognizable to humans; (3) their robustness varies substantially across datasets and harm categories, with multimodal content and self-harm exhibiting pronounced vulnerabilities. This yields the conclusion that replacing conventional moderation classifiers with foundation-model-based APIs does not, by itself, provide a reliable security boundary. Such systems must be evaluated under realistic transformations and deployed as one component of a layered moderation pipeline rather than as standalone safety filters.

内容审核模型鲁棒性安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。