arXiv:2503.09964cs.CRcs.CL2025-03EMNLP被引 1

测试大模型对AI生成极端内容的防御能力,发现现有安全机制形同虚设。

ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content

  • 用最新图像生成技术构造逼真极端内容样本
  • 顶尖模型在攻击下仍以高成功率生成极端内容
  • 适合安全研究者与模型开发者参考

大型多模态模型(LMMs)正日益面临由AI生成的极端内容威胁,包括逼真的图像和文本,可绕过安全机制并生成有害输出。然而,现有评估数据集对极端内容的探索有限,普遍缺乏AI生成图像、多样化的图像生成模型以及对历史事件的全面覆盖,阻碍了对模型漏洞的完整评估。为此,我们提出ExtremeAIGC,一个用于评估LMM对极端内容脆弱性的基准数据集与评估框架。该数据集通过采用最先进的图像生成技术,构建涵盖多种文本与图像样本的多样化案例,模拟真实世界事件与恶意使用场景。我们的研究揭示了令人震惊的模型弱点,表明即使最前沿的安全措施也无法有效阻止极端内容的生成。我们系统量化了各类攻击策略的成功率,暴露了当前防御体系的关键缺陷,强调亟需更强大的缓解策略。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) are increasingly vulnerable to AI-generated extremist content, including photorealistic images and text, which can be used to bypass safety mechanisms and generate harmful outputs. However, existing datasets for evaluating LMM robustness offer limited exploration of extremist content, often lacking AI-generated images, diverse image generation models, and comprehensive coverage of historical events, which hinders a complete assessment of model vulnerabilities. To fill this gap, we introduce ExtremeAIGC, a benchmark dataset and evaluation framework designed to assess LMM vulnerabilities against such content. ExtremeAIGC simulates real-world events and malicious use cases by curating diverse text- and image-based examples crafted using state-of-the-art image generation techniques. Our study reveals alarming weaknesses in LMMs, demonstrating that even cutting-edge safety measures fail to prevent the generation of extremist material. We systematically quantify the success rates of various attack strategies, exposing critical gaps in current defenses and emphasizing the need for more robust mitigation strategies.

多模态安全极端内容模型评测AI风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。