arXiv:2606.23267cs.CVcs.CY2026-06

提出无需训练的快速安全生成方法,4步内精准移除敏感内容

Safe Few-Step Generation via Velocity Editing

论文配图:Safe Few-Step Generation via Velocity Editing
图 1 · 摘自论文原文
  • 通过编辑速度场实现安全生成,不改变原始提示词
  • 4步生成下使攻击成功率降至6.3%~6.8%,保持良好图像质量
  • 适合对生成安全性和效率要求高的实际应用

流匹配(Flow matching)已成为最先进的文本到图像(T2I)生成范式,可在极少数采样步骤内生成高质量图像。随着其在真实场景中的广泛应用,确保生成内容的安全性成为关键挑战。现有方法多依赖迭代轨迹调整或基于CLIP的提示嵌入操作,但在流匹配框架下受限于采样步数少、上下文感知编码器强,效果不佳。本文提出VESFlow,一种无需训练的安全方法,专为极小步数的流匹配设计。利用流匹配模型学习的边际速度特性,通过安全条件后验直接编辑速度场,引导生成安全结果而不修改原始提示。基于该方法在良性提示下输出不变的特性,进一步引入基于风险评分的过滤机制,避免不必要的计算。在此基础上提出VESFlow+,不仅将速度推向安全方向,还主动远离危险方向。实验表明,在4步均值流模型上,VESFlow+可将NudeNet对Ring-A-Bell和MMA-Diffusion的攻击成功率分别降低至6.3%和6.8%,同时保持良性提示下的生成保真度。

原文摘要 · Abstract (English)

Flow matching has recently emerged as a strong paradigm for state-of-the-art text-to-image (T2I) generation, enabling high-quality generation with a small number of sampling steps. As these models are increasingly integrated into real-world applications, ensuring safe and non-sensitive content generation has become a critical requirement. However, adapting safety and concept removal methods to this new generation framework remains an open challenge. Specifically, prior methods largely rely on iterative trajectory steering across a number of denoising steps or on CLIP-centric prompt embedding manipulation. These design assumptions pose fundamental bottlenecks for safety in flow matching-based T2I generation, where limited sampling steps constrain iterative correction and modern context-aware text encoders diminish the effectiveness of embedding-level interventions. In this paper, we propose VESFlow, a training-free safety method tailored to flow matching with extremely few sampling steps. Leveraging the fact that flow matching models learn the marginal velocity, we directly edit the velocity field via a safe-conditional posterior. VESFlow steers the trajectory toward safe outputs while leaving the conditioning prompt unchanged. Building on the observation that VESFlow leaves outputs unchanged under benign prompts, we further introduce a risk score-based filtering that bypasses velocity editing to reduce computational cost while preserving benign prompt generation. Based on this filtering, we propose VESFlow+, a stronger variant of VESFlow that not only edits the velocity toward the safe direction, but also pushes it away from the unsafe direction. Experimental results show that VESFlow+ removes the target concept, reducing the attack success rate by NudeNet to 6.3% on Ring-A-Bell and 6.8% on MMA-Diffusion on the 4-step MeanFlow model, while preserving fidelity on benign prompts.

文本生成安全生成扩散模型速度编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。