arXiv:2606.09901cs.GRcs.CV2026-06

揭示扩散模型图像编辑中可控性与保真度的权衡边界

On the Controllability-Fidelity Frontier in Diffusion Editing

  • 从数学上分析噪声注入与梯度引导的动态机制
  • 发现身份漂移、提示敏感等关键失败模式
  • 适合关注图像编辑可靠性与伦理风险的研究者

基于扩散的生成模型虽具备强大图像编辑能力,但实现精确控制同时保持内容保真与安全仍具挑战。本文对可控扩散编辑进行理论与实证研究,分析用户意图遵循性、非目标内容保留与输出质量间的权衡。涵盖文本与掩码引导编辑、点拖拽操作及反演流程。推导编辑目标的数学表达,分析噪声注入、得分引导与反演误差的动力学特性。给出重建误差、重复编辑稳定性及变化局部性的理论界。提出掩码局部化与指令引导编辑的算法框架(附伪代码),并在多任务与指标(FID、身份相似度、CLIP对齐度、伪影评分等)下对比SOTA方法(如TF-ICON、DragFlow、InstructPix2Pix、UltraEdit)。结果揭示身份漂移、提示敏感性与组合错误等关键失败模式。讨论滥用风险、偏见、同意问题及概念擦除技术(如MACE、ANT、EraseAnything)作为防护措施。最后提出负责任高保真编辑的最佳实践与未来方向。

原文摘要 · Abstract (English)

Diffusion-based generative models enable powerful image editing capabilities, but achieving precise control while maintaining fidelity and safety remains challenging. We present a comprehensive theoretical and empirical study of controllable diffusion-based image editing, analyzing the trade-offs between adherence to user intent, preservation of non-target content, and output quality. Our work spans text- and mask-guided edits, point/drag manipulation, and inversion-based pipelines. We derive mathematical formulations of editing objectives and analyze dynamics of noise injection, score guidance, and inversion error. We provide theoretical bounds on reconstruction error, stability under repeated edits, and locality of changes. We propose algorithmic frameworks (with pseudocode) for mask-localized and instruction-guided editing, and present extensive experiments comparing state-of-the-art methods (e.g.\ TF-ICON \cite{lu2023tficone}, DragFlow \cite{zhou2025dragflow}, InstructPix2Pix \cite{brooks2023instructpix2pix}, UltraEdit \cite{zhao2024ultraedit}) on multiple tasks and metrics (FID, identity similarity, CLIP alignment, artifact scores, etc). Our results reveal key failure modes, such as identity drift, prompt sensitivity, and compositional errors. We also discuss ethical considerations in image editing, including misuse risks, bias, consent, and concept erasure techniques (e.g.\ MACE \cite{lu2024mace}, ANT \cite{li2025ant}, EraseAnything \cite{gao2024eraseanything}) as safeguards. We conclude with best practices and future directions for responsible, high-fidelity diffusion-based editing.

扩散模型图像编辑可控性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。