用多轮编辑提升文生图模型安全性,兼顾安全与可用性。
SafeEditor: Unified MLLM for Efficient Post-hoc T2I Safety Editing
- 模仿人类认知,通过多轮交互式编辑修复生成图像中的安全问题。
- 在多个数据集上实现更低的过度拒绝率,安全与实用平衡更优。
- 无需重训练,可适配任意文生图模型,部署灵活高效。
随着文本到图像(T2I)模型的快速发展,其安全性日益重要。现有方法分为训练时和推理时两类,其中推理时方法因成本低被广泛采用,但常存在过度拒绝和安全-效用失衡的问题。为此,我们提出一种多轮安全编辑框架,作为模型无关、即插即用的模块,可高效对任意T2I模型进行安全对齐。核心是为安全编辑构建的多轮图文交错数据集MR-SafeEdit。我们提出后置安全编辑范式,模拟人类识别并修正不安全内容的过程。基于此,开发出SafeEditor——一个统一的多模态大模型(MLLM),支持对生成图像进行多轮安全编辑。实验表明,SafeEditor在降低过度拒绝的同时,实现了更优的安全-效用平衡,优于现有方法。
原文摘要 · Abstract (English)
With the rapid advancement of text-to-image (T2I) models, ensuring their safety has become increasingly critical. Existing safety approaches can be categorized into training-time and inference-time methods. While inference-time methods are widely adopted due to their cost-effectiveness, they often suffer from limitations such as over-refusal and imbalance between safety and utility. To address these challenges, we propose a multi-round safety editing framework that functions as a model-agnostic, plug-and-play module, enabling efficient safety alignment for any text-to-image model. Central to this framework is MR-SafeEdit, a multi-round image-text interleaved dataset specifically constructed for safety editing in text-to-image generation. We introduce a post-hoc safety editing paradigm that mirrors the human cognitive process of identifying and refining unsafe content. To instantiate this paradigm, we develop SafeEditor, a unified MLLM capable of multi-round safety editing on generated images. Experimental results show that SafeEditor surpasses prior safety approaches by reducing over-refusal while achieving a more favorable safety-utility balance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。