不修改模型也能生成安全图像,还能保持原图结构不变。
Safety Without Semantic Disruptions: Editing-free Safe Image Generation via Context-preserving Dual Latent Reconstruction
- 用双隐空间重构技术,在不破坏语义结构的前提下过滤有害内容。
- 在多个安全生成基准上达到当前最优表现,且支持安全等级灵活调节。
- 适合关注生成内容安全性的研究人员与应用开发者使用。
在大规模未清洗数据上训练多模态生成模型可能导致用户接触到有害、不安全或文化不当的内容。尽管已有模型编辑方法用于移除嵌入空间和隐空间中的不良概念,但可能无意中破坏学习到的流形结构,导致邻近语义概念发生扭曲。我们揭示了现有模型编辑技术的局限性,表明即使良性且接近的概念也可能出现错位。为解决安全生成需求,我们采用安全嵌入并引入可调加权求和的改进扩散过程,在隐空间中实现更安全的图像生成。该方法在不损害学习流形结构的前提下保留全局上下文。我们在安全图像生成基准上取得当前最优结果,并提供直观的安全等级控制。我们识别出安全性与审查之间的权衡关系,为伦理AI模型的发展提供了必要视角。代码将公开。
原文摘要 · Abstract (English)
Training multimodal generative models on large, uncurated datasets can result in users being exposed to harmful, unsafe and controversial or culturally-inappropriate outputs. While model editing has been proposed to remove or filter undesirable concepts in embedding and latent spaces, it can inadvertently damage learned manifolds, distorting concepts in close semantic proximity. We identify limitations in current model editing techniques, showing that even benign, proximal concepts may become misaligned. To address the need for safe content generation, we leverage safe embeddings and a modified diffusion process with tunable weighted summation in the latent space to generate safer images. Our method preserves global context without compromising the structural integrity of the learned manifolds. We achieve state-of-the-art results on safe image generation benchmarks and offer intuitive control over the level of model safety. We identify trade-offs between safety and censorship, which presents a necessary perspective in the development of ethical AI models. We will release our code. Keywords: Text-to-Image Models, Generative AI, Safety, Reliability, Model Editing
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。