用人类创意+机器扩展,生成更真实多样的文本生成模型攻击提示。
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models
- 结合人类创意与机器能力,扩增多样化的攻击提示种子。
- 新数据集覆盖535个地理区域,熵值达7.48,显著提升多样性。
- 适合关注AI安全性、对抗样本生成的研究者与开发者。
文本到图像(T2I)模型在众多应用中日益普及,对其抵御对抗性攻击的稳健评估成为关键任务。持续获取跨领域的新颖且具有挑战性的对抗性提示对压力测试模型抵御多路径新型攻击至关重要。当前生成此类提示的方法要么完全依赖人工创作,要么由机器合成。前者数据量小且文化语境分布不均;后者虽可规模化,但缺乏人类提示的真实细节与创造性攻击策略。为此,我们提出Seed2Harvest,一种混合式红队测试方法,通过有指导地扩展文化多元的人类创作攻击提示种子。生成的提示保留了人类提示的特征与攻击模式,同时保持相当的平均攻击成功率:NudeNet为0.31,SD NSFW为0.36,Q16为0.12。扩充后的数据集在多样性上显著提升,包含535个独特地理位置,香农熵达7.48,相较原数据集的58个位置与5.28熵值大幅提升。本工作表明,人机协作能有效结合人类创造力与机器计算能力,实现全面且可扩展的T2I模型安全持续评估。
原文摘要 · Abstract (English)
Text-to-image (T2I) models have become prevalent across numerous applications, making their robust evaluation against adversarial attacks a critical priority. Continuous access to new and challenging adversarial prompts across diverse domains is essential for stress-testing these models for resilience against novel attacks from multiple vectors. Current techniques for generating such prompts are either entirely authored by humans or synthetically generated. On the one hand, datasets of human-crafted adversarial prompts are often too small in size and imbalanced in their cultural and contextual representation. On the other hand, datasets of synthetically-generated prompts achieve scale, but typically lack the realistic nuances and creative adversarial strategies found in human-crafted prompts. To combine the strengths of both human and machine approaches, we propose Seed2Harvest, a hybrid red-teaming method for guided expansion of culturally diverse, human-crafted adversarial prompt seeds. The resulting prompts preserve the characteristics and attack patterns of human prompts while maintaining comparable average attack success rates (0.31 NudeNet, 0.36 SD NSFW, 0.12 Q16). Our expanded dataset achieves substantially higher diversity with 535 unique geographic locations and a Shannon entropy of 7.48, compared to 58 locations and 5.28 entropy in the original dataset. Our work demonstrates the importance of human-machine collaboration in leveraging human creativity and machine computational capacity to achieve comprehensive, scalable red-teaming for continuous T2I model safety evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。