通过认知变形攻击,让AI图像生成模型在保留主体的前提下植入有害情境。
CogMorph: Cognitive Morphing Attacks for Text-to-Image Models
- 基于人类认知原理,分层提取原提示关键元素并注入毒性特征。
- 构建1176个高质量有毒提示,使攻击成功率平均提升20.62%。
- 揭示文本生成图像模型的隐性伦理风险,适合安全与伦理研究者参考。
文本到图像(T2I)生成模型虽推动了创意设计发展,但本文揭示其存在此前未被识别的伦理风险,并提出认知变形攻击(CogMorph)方法。该方法在保留原始核心主体的基础上,嵌入有毒或有害的上下文元素,利用人类对概念的认知依赖整体视觉场景的特性,显著放大情感伤害。为此,研究构建涵盖10大类、48子类的图像毒性分类体系,并形成包含1176个高质量有毒提示的毒性风险矩阵。CogMorph引入认知毒性增强技术,建立富含人类外部毒性表征(如细粒度视觉特征)的知识库,指导对抗性提示优化;同时提出上下文分层变形机制,分层提取原提示中的场景、主体、身体部位等关键成分,迭代检索并融合毒性特征以实现有害上下文注入。在多个开源T2I模型及黑盒商业API(如DALL·E 3)上的实验表明,该方法显著优于现有基线,平均性能提升20.62%。
原文摘要 · Abstract (English)
The development of text-to-image (T2I) generative models, that enable the creation of high-quality synthetic images from textual prompts, has opened new frontiers in creative design and content generation. However, this paper reveals a significant and previously unrecognized ethical risk inherent in this technology and introduces a novel method, termed the Cognitive Morphing Attack (CogMorph), which manipulates T2I models to generate images that retain the original core subjects but embeds toxic or harmful contextual elements. This nuanced manipulation exploits the cognitive principle that human perception of concepts is shaped by the entire visual scene and its context, producing images that amplify emotional harm far beyond attacks that merely preserve the original semantics. To address this, we first construct an imagery toxicity taxonomy spanning 10 major and 48 sub-categories, aligned with human cognitive-perceptual dimensions, and further build a toxicity risk matrix resulting in 1,176 high-quality T2I toxic prompts. Based on this, our CogMorph first introduces Cognitive Toxicity Augmentation, which develops a cognitive toxicity knowledge base with rich external toxic representations for humans (e.g., fine-grained visual features) that can be utilized to further guide the optimization of adversarial prompts. In addition, we present Contextual Hierarchical Morphing, which hierarchically extracts critical parts of the original prompt (e.g., scenes, subjects, and body parts), and then iteratively retrieves and fuses toxic features to inject harmful contexts. Extensive experiments on multiple open-sourced T2I models and black-box commercial APIs (e.g., DALLE-3) demonstrate the efficacy of CogMorph which significantly outperforms other baselines by large margins (+20.62% on average).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。