测试大模型图像生成安全漏洞,设计动态基准评估攻击风险。
Unmasking the Canvas: A Dynamic Benchmark for Image Generation Jailbreaking and LLM Content Safety
- 用结构化提示与多语言伪装构造攻击样本。
- 在LLaMA-3上验证出多种真实伪造图像生成案例。
- 提供可演化、分级标注的动态评测数据集,适合安全研究者使用。
当前大型语言模型在图像生成任务中表现卓越,但内容安全机制仍易受提示攻击。通过在ChatGPT、MetaAI和Grok等平台初步测试发现,简短自然的提示即可引发生成包含伪造文件或公众人物被篡改的敏感图像。为此,我们提出动态可扩展的评测基准Unmasking the Canvas(UTCB),结合结构化提示工程、多语言混淆(如祖鲁语、盖尔语、Base64)及基于Groq部署的LLaMA-3进行评估。该流程支持零样本与备用提示策略、风险评分与自动标记,所有生成结果均附带丰富元数据,并按可信度分为青铜(未验证)、白银(模型辅助验证)、黄金(人工验证)三级。UTCB将随新数据源、提示模板与模型行为持续演进。注意:本文含用于测试安全性的对抗性输入视觉示例,所有输出已脱敏以确保负责任披露。
原文摘要 · Abstract (English)
Existing large language models (LLMs) are advancing rapidly and produce outstanding results in image generation tasks, yet their content safety checks remain vulnerable to prompt-based jailbreaks. Through preliminary testing on platforms such as ChatGPT, MetaAI, and Grok, we observed that even short, natural prompts could lead to the generation of compromising images ranging from realistic depictions of forged documents to manipulated images of public figures. We introduce Unmasking the Canvas (UTC Benchmark; UTCB), a dynamic and scalable benchmark dataset to evaluate LLM vulnerability in image generation. Our methodology combines structured prompt engineering, multilingual obfuscation (e.g., Zulu, Gaelic, Base64), and evaluation using Groq-hosted LLaMA-3. The pipeline supports both zero-shot and fallback prompting strategies, risk scoring, and automated tagging. All generations are stored with rich metadata and curated into Bronze (non-verified), Silver (LLM-aided verification), and Gold (manually verified) tiers. UTCB is designed to evolve over time with new data sources, prompt templates, and model behaviors. Warning: This paper includes visual examples of adversarial inputs designed to test model safety. All outputs have been redacted to ensure responsible disclosure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。