多模态大模型理解力强,却更易生成危险内容且难被检测。
When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigm
- 利用文本理解能力生成复杂指令下的违规图像
- 在多个基准测试中,生成的不安全图像比例高于扩散模型
- 现有检测器难以识别其生成内容,长描述输入可绕过防护
近期,多模态大语言模型(MLLMs)作为统一的语言与图像生成范式兴起。相比扩散模型,MLLMs 具有更强的语义理解能力,能处理更复杂的文本输入并理解更丰富的上下文含义。然而,这种增强的理解能力也可能引入新的、潜在更大的安全风险。以扩散模型为参照,我们从两个维度系统分析并比较了新兴 MLLMs 的安全风险:不安全内容生成与虚假图像合成。在多个不安全生成基准数据集上,我们发现 MLLMs 生成的不安全图像数量显著高于扩散模型。这一差异部分源于扩散模型常无法理解抽象提示,产生损坏输出,而 MLLMs 能理解这些提示并生成违规内容。对于当前先进的假图像检测器,MLLM 生成的图像也更难被识别。即使检测器使用 MLLM 特定数据重新训练,仍可通过提供更长、更详细的输入实现绕过。测量结果显示,前沿生成范式 MLLMs 的安全风险尚未被充分认识,给现实世界安全带来新挑战。
原文摘要 · Abstract (English)
Recently, multimodal large language models (MLLMs) have emerged as a unified paradigm for language and image generation. Compared with diffusion models, MLLMs possess a much stronger capability for semantic understanding, enabling them to process more complex textual inputs and comprehend richer contextual meanings. However, this enhanced semantic ability may also introduce new and potentially greater safety risks. Taking diffusion models as a reference point, we systematically analyze and compare the safety risks of emerging MLLMs along two dimensions: unsafe content generation and fake image synthesis. Across multiple unsafe generation benchmark datasets, we observe that MLLMs tend to generate more unsafe images than diffusion models. This difference partly arises because diffusion models often fail to interpret abstract prompts, producing corrupted outputs, whereas MLLMs can comprehend these prompts and generate unsafe content. For current advanced fake image detectors, MLLM-generated images are also notably harder to identify. Even when detectors are retrained with MLLMs-specific data, they can still be bypassed by simply providing MLLMs with longer and more descriptive inputs. Our measurements indicate that the emerging safety risks of the cutting-edge generative paradigm, MLLMs, have not been sufficiently recognized, posing new challenges to real-world safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。