arXiv:2511.11693cs.AIcs.CR2025-11AAAI被引 2

用AI自动重写危险提示,让图像生成既安全又不失创意。

Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation

论文配图:Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation
图 1 · 摘自论文原文
  • 通过多层检测识别内容风险、文化规范与隐含意图
  • 零样本下重写提示,安全率提升最高达100%
  • 适合需要高安全性的开放场景图像生成应用

生成式视觉语言模型如Stable Diffusion在创意媒体合成方面表现卓越,但面对恶意提示时可能生成不安全、冒犯性或文化不当的内容。现有防御手段难以在不牺牲生成质量或增加成本的前提下实现人类价值观对齐。为此,我们提出VALOR(Value-Aligned LLM-Overseen Rewriter),一个模块化、零样本的代理式框架,用于更安全且更有帮助的文本到图像生成。VALOR结合分层提示分析与人类价值推理:多级NSFW检测器过滤词汇与语义风险;文化价值对齐模块识别社会规范、合法性及代表性伦理违规;意图消歧器探测微妙或间接的不安全含义。一旦检测到不安全内容,提示将由大语言模型在动态角色指令下选择性重写,以保留用户意图同时确保对齐。若生成图像仍不通过安全检查,VALOR可选进行风格重生成,将输出导向更安全的视觉领域而不改变核心语义。在对抗性、模糊性和价值敏感性提示上的实验表明,VALOR将不安全输出减少最多达100.00%,同时保持提示有用性和创造力。结果表明,VALOR是开放世界中部署安全、对齐且有益图像生成系统的可扩展有效方法。

原文摘要 · Abstract (English)

Generative vision-language models like Stable Diffusion demonstrate remarkable capabilities in creative media synthesis, but they also pose substantial risks of producing unsafe, offensive, or culturally inappropriate content when prompted adversarially. Current defenses struggle to align outputs with human values without sacrificing generation quality or incurring high costs. To address these challenges, we introduce VALOR (Value-Aligned LLM-Overseen Rewriter), a modular, zero-shot agentic framework for safer and more helpful text-to-image generation. VALOR integrates layered prompt analysis with human-aligned value reasoning: a multi-level NSFW detector filters lexical and semantic risks; a cultural value alignment module identifies violations of social norms, legality, and representational ethics; and an intention disambiguator detects subtle or indirect unsafe implications. When unsafe content is detected, prompts are selectively rewritten by a large language model under dynamic, role-specific instructions designed to preserve user intent while enforcing alignment. If the generated image still fails a safety check, VALOR optionally performs a stylistic regeneration to steer the output toward a safer visual domain without altering core semantics. Experiments across adversarial, ambiguous, and value-sensitive prompts show that VALOR significantly reduces unsafe outputs by up to 100.00% while preserving prompt usefulness and creativity. These results highlight VALOR as a scalable and effective approach for deploying safe, aligned, and helpful image generation systems in open-world settings.

图像生成安全对齐提示重写零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。